LLMpediaThe first transparent, open encyclopedia generated by LLMs

Literary and Linguistic Computing

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Literary and Linguistic Computing
NameLiterary and Linguistic Computing
DisciplineComputational humanities
SubdisciplineDigital philology, corpus linguistics, text analysis

Literary and Linguistic Computing Literary and Linguistic Computing is an interdisciplinary field connecting computational methods with William Shakespeare, Geoffrey Chaucer, Jane Austen, James Joyce, and Homer studies, influencing work at institutions such as University of Oxford, University of Cambridge, Harvard University, Stanford University, and University College London; it engages scholars associated with projects like the Oxford English Dictionary, the Google Books project, the Perseus Digital Library, the Project Gutenberg, and the Text Encoding Initiative. The field intersects with scholarship on figures such as T. S. Eliot, Virginia Woolf, Leo Tolstoy, Fyodor Dostoevsky, and Marcel Proust, and with centers including the British Library, the Library of Congress, the Bibliothèque nationale de France, the Max Planck Institute, and the Institute for Advanced Study.

History

The discipline traces roots to early computing experiments linked to people like Alan Turing, John von Neumann, H. P. Lovecraft-era text collections, and initiatives such as the Index Thomisticus, the CELEX lexical database, and the Brown Corpus, while later milestones include connections to projects at University of Birmingham, King's College London, Yale University, Columbia University, and Princeton University. Early methodological shifts were influenced by debates in forums attended by scholars from MIT, Bell Labs, IBM, RAND Corporation, and SRI International, and by conferences convened at ACL Anthology venues, the European Association for Digital Humanities meetings, and symposia at the American Council of Learned Societies. Developments in textual encoding and standards emerged alongside initiatives like the Text Encoding Initiative, collaborations with the Oxford English Dictionary editorial teams, and funding from agencies such as the National Endowment for the Humanities, the European Research Council, and the Wellcome Trust.

Scope and Methods

The scope covers computational philology applied to corpora involving figures such as Charles Dickens, Emily Dickinson, Mark Twain, Pablo Neruda, and Li Bai, employing methods drawn from Noam Chomsky-inspired linguistics, Ferdinand de Saussure-informed structuralism, Roman Jakobson's poetics, and techniques used in projects at Google Research, Microsoft Research, Facebook AI Research, OpenAI, and DeepMind. Core methods include corpus annotation practices developed in the Brown Corpus, concordancing techniques used by scholars at University of Sheffield and Lancaster University, statistical models influenced by work at Bell Labs and AT&T, machine learning routines derived from research at Carnegie Mellon University and University of Toronto, and digital editing protocols shaped by the Bodleian Library and Bibliothèque nationale de France.

Tools and Technologies

Practitioners use software and platforms including Python (programming language), R (programming language), Perl, TEI (Text Encoding Initiative), XML, Unicode, MySQL, PostgreSQL, Apache Hadoop, Apache Spark, TensorFlow, PyTorch, NLTK, spaCy, and interfaces developed at labs like Massachusetts Institute of Technology and Stanford University. Text digitization workflows rely on equipment and standards from Zanichelli-era scanners, OCR engines popularized by ABBYY, data curation practices from Digital Humanities Observatory, and repository models used by Europeana, the HathiTrust, the Internet Archive, and the National Library of Scotland.

Applications and Case Studies

Applications span stylometric studies of authorship attribution for texts by Edward Said, Roland Barthes, Michel Foucault, Hannah Arendt, and J. R. R. Tolkien; computational editions of manuscripts associated with Dante Alighieri, Geoffrey Chaucer, John Milton, Miguel de Cervantes, and Gilgamesh fragments; and sociolinguistic corpora analysis in projects linked to UNESCO, World Bank, European Commission, British Library, and Library of Congress. Case studies include digital reconstructions of works related to Sappho, computational reception studies of Oscar Wilde and Marcel Proust, network analysis of correspondence involving Jane Austen and Samuel Johnson, and topic-modeling investigations into periodicals such as The Times, The New Yorker, and Le Monde.

Education and Training

Training pathways exist in programs at University of Oxford, University of Cambridge, King's College London, University of Edinburgh, University of Leeds, Columbia University, New York University, University of California, Berkeley, University of Toronto, and Australian National University, often drawing on coursework in computational methods from Massachusetts Institute of Technology, Stanford University, Princeton University, Yale University, and summer schools sponsored by Digital Humanities Summer Institute. Degrees and certificates connect to departments such as Faculty of History, University of Oxford, School of Advanced Study, University of London, and centers like the Royal Historical Society and training funded through grants from the Andrew W. Mellon Foundation.

Criticism and Ethical Issues

Critiques reference debates familiar to scholars of Michel Foucault, Pierre Bourdieu, Jürgen Habermas, Edward Said, and Gayatri Chakravorty Spivak about representation, bias, and authority in textual corpora, and raise concerns echoed in discussions at United Nations forums, panels at the AAAI Conference, and statements by organizations such as ACM and IEEE about algorithmic transparency. Ethical issues involve provenance debates exemplified by controversies at the British Museum, disputed digitization cases involving the National Archives (UK), questions of consent addressed by the World Health Organization, and policy considerations in line with guidelines from the European Commission and the National Institutes of Health.

Category:Digital humanities