This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| Corpus del Español (Mark Davies) | |
|---|---|
| Name | Corpus del Español |
| Creator | Mark Davies |
| Type | Historical and contemporary Spanish corpus |
| Languages | Spanish |
| Release | 2000s |
| Access | Subscription / institutional |
| License | Proprietary / academic |
Corpus del Español (Mark Davies) is a large-scale, annotated corpus of Spanish created and maintained by linguist Mark Davies. The corpus aggregates historical and contemporary texts to support research in lexicography, historical linguistics, philology, and computational linguistics, and it has been used by scholars associated with institutions across Spain, Latin America, and the United States. It connects textual evidence from diverse sources such as archives, libraries, publishers, and universities to enable empirical study of Spanish across centuries.
The project was developed by Mark Davies at Brigham Young University with collaborations involving Real Academia Española, Universidad de Salamanca, Biblioteca Nacional de España, Universidad Complutense de Madrid, and research centers like Centro de Estudios Históricos and Instituto Cervantes. The corpus includes materials spanning from the medieval period to the present, drawing on texts from collections such as Biblioteca Vaticana, Archivo General de Indias, Archivo Histórico Nacional, British Library, and Library of Congress. Users include scholars affiliated with Universidad Autónoma de Madrid, University of Oxford, Harvard University, Stanford University, Yale University, and Universidad de Buenos Aires.
The project originated in the late 1990s and expanded through collaborations with institutions including National Endowment for the Humanities, European Commission, Wellcome Trust, and Pontificia Universidad Católica de Chile. Early development drew on digitization projects at Hispanic Society of America and Fundación Biblioteca Virtual Miguel de Cervantes, and incorporated editorial standards influenced by scholars from Real Academia Española and departments at University of California, Berkeley and University of Texas at Austin. Funding and technical partnerships involved organizations like JSTOR, Google Books project partners, and university presses including Oxford University Press and Cambridge University Press for comparative corpus work. The corpus evolved alongside projects such as the Corpus of Contemporary American English and the Frantext database, while being cited in studies using methods from Noam Chomsky, William Labov, and Ferdinand de Saussure-inspired traditions.
The compilation integrates multiple subcorpora: a historical Spanish component with texts from authors such as Miguel de Cervantes, Lope de Vega, Garcilaso de la Vega, Fray Luis de León, Gustavo Adolfo Bécquer, and Benito Pérez Galdós; a modern written component featuring newspapers and magazines like El País, ABC, La Vanguardia, El Universal, and Clarín; and a spoken component with recordings and transcripts from institutions including Radio Nacional de España, RTVE, BBC Mundo, Televisión Española, and university corpora from Universidad Nacional Autónoma de México. It draws on legal texts from archives such as Boletín Oficial del Estado and literary works archived at Bibliothèque nationale de France and Biblioteca Nacional de Chile. The collection also references texts associated with figures like Sor Juana Inés de la Cruz, Jorge Luis Borges, Pablo Neruda, Gabriel García Márquez, Federico García Lorca, Julio Cortázar, Mario Vargas Llosa, and Isabel Allende.
Annotation conventions follow standards influenced by projects from TEI Consortium, ISO, and research groups at Universitat Pompeu Fabra and Consejo Superior de Investigaciones Científicas. The corpus uses tokenization, part-of-speech tagging, and lemmatization protocols comparable to those deployed by Penn Treebank initiatives and computational linguistics labs at Massachusetts Institute of Technology and Carnegie Mellon University. Corpus annotation incorporates guidelines from projects involving Mark Aronoff-style morphology, statistical methods akin to work by Geoffrey Hinton and Yoshua Bengio in machine learning for language, and evaluation practices used by groups at Google Research and Microsoft Research. Metadata schemas align with cataloging standards from Dublin Core and cooperation with libraries such as Biblioteca Nacional de España and British Library.
Access is provided via an online interface hosted by Brigham Young University with subscription options tailored for academic institutions, libraries like Harvard Library and Biblioteca Nacional de España, and commercial partners including ProQuest and EBSCO Information Services. The web interface affords concordance searches, frequency lists, and collocation tools inspired by interfaces from Sketch Engine and AntConc. Licensing models resemble those of digital resources distributed by Gale (publisher), Elsevier, and university presses, with negotiated terms for corpus mining by researchers from University College London and Universidad de Sevilla.
The corpus has informed studies in historical lexicography, cited in works from scholars at Universidad de Salamanca, University of Chicago, Columbia University, University of Pennsylvania, University of Toronto, and University of Michigan. It has been used in projects on diachronic change relevant to researchers associated with Real Academia Española and in computational studies conducted at Stanford NLP Group, Barcelona Supercomputing Center, and Max Planck Institute for Psycholinguistics. The resource underpins dictionaries, concordances, and reference works produced by Oxford University Press, and has been employed in dissertations and monographs at institutions like Universidad de Granada and Universidad Complutense de Madrid.
Critiques have focused on representativeness and licensing constraints noted by scholars at Universidad de Santiago de Compostela, Universidad de La Plata, and independent researchers associated with Open Knowledge Foundation. Concerns mirror debates seen in relation to Google Books regarding copyright, selection bias highlighted in work from JSTOR critics, and challenges of OCR accuracy noted by librarians at Biblioteca Nacional de España and technical teams at British Library. Methodological limitations referenced by computational linguists at University of Edinburgh and ETH Zurich include uneven diachronic coverage, metadata inconsistencies, and proprietary access limiting reproducibility compared to open datasets promoted by Common Crawl advocates.
Category:Corpora