LLMpediaThe first transparent, open encyclopedia generated by LLMs

Mark Davies

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

Mark Davies
NameMark Davies
Birth date1957
Birth placeLeeds
OccupationCorpus linguist, lexicographer, professor
Alma materUniversity of Cambridge, University of Oxford
Known forCorpus linguistics, word frequency databases, language corpora

Mark Davies is a British-born corpus linguist and lexicographer noted for pioneering large-scale electronic corpora, corpus-driven lexicography, and frequency-based language description. He established several widely used corpora and analytical tools that have influenced research at institutions such as University of Birmingham, Brigham Young University, and national projects connected to British National Corpus initiatives. His work has intersected with projects in computational linguistics, digital humanities, and language technology across the United Kingdom, the United States, and beyond.

Early life and education

Born in Leeds in 1957, Davies completed his undergraduate studies at University of Cambridge where he read linguistics and English language. He pursued graduate research at University of Oxford with a focus on corpus-based approaches influenced by earlier work at Lancaster University and the development of the Oxford English Dictionary's corpus resources. During postgraduate training he engaged with researchers at University of Birmingham and collaborators at COBUILD and the British National Corpus project, situating his early career within a network that included scholars from University of Nottingham, University College London, and University of Manchester.

Academic and professional career

Davies began his academic appointment in the 1980s and progressed through roles including lecturer, senior lecturer, and professor at institutions such as Brigham Young University and University of Birmingham. He founded and directed major corpus projects and digital archives hosted at university centers akin to those at Corpus linguistics research group environments. His technical contributions included the design and implementation of web-accessible concordancers and frequency lists that paralleled tools developed at COBUILD, Lancaster University Centre for Corpus Research, and the International Computer Archive of Modern and Medieval English (ICAME). He collaborated with computing groups at Stanford University and Massachusetts Institute of Technology on text processing pipelines and worked with publishing houses such as Oxford University Press and Cambridge University Press on corpus-informed lexicography. Over his career he supervised graduate students who took positions at universities including University of California, Los Angeles, University of Texas at Austin, University of Sydney, and University of Toronto.

Major contributions and publications

Davies is best known for developing scalable corpora and making them publicly accessible through user-friendly interfaces. He created large web-based corpora that provided frequency lists, collocation data, and concordance views used by scholars at Harvard University, Yale University, University of Edinburgh, and research centers in Germany, China, and Japan. His published work spans journal articles in outlets associated with Journal of English Linguistics, Applied Linguistics, and proceedings of conferences such as Computational Linguistics (ACL) and Corpus Linguistics (ICAME) meetings. Key publications include introductions to corpus methodology used alongside textbooks from Cambridge University Press and methodological papers often cited by projects at Google Research, Microsoft Research, and national language institutes like the Institute for Language and Speech Processing. He compiled and released specialized corpora for varieties of English—American, British, Irish, Australian—that have been integrated into research at National Institute for Japanese Language and Linguistics and comparative studies at European Research Council projects. His work on frequency effects, lexical change, and register variation influenced subsequent monographs from authors affiliated with Columbia University, University of Pennsylvania, and Princeton University.

Awards and honors

Davies received recognition from professional bodies including awards from European Association for Corpus Linguistics and fellowships linked to centers such as British Academy and national research councils comparable to National Science Foundation. He was invited to deliver keynote lectures at major venues including American Association for Applied Linguistics conferences, plenaries at International Corpus Linguistics Conference, and symposia hosted by Royal Society of Edinburgh. Honorary appointments and visiting professorships took him to institutions such as University of Hong Kong, University of Cape Town, and University of Melbourne. His datasets and tools have been lauded in citation indices maintained by Web of Science and Google Scholar for their broad uptake across humanities and computational fields.

Personal life and legacy

Outside academia, Davies has been active in initiatives advocating open data and reproducible research, engaging with organizations like Open Knowledge Foundation and community efforts similar to Wikimedia Foundation projects. Colleagues and former students often cite his emphasis on empirical rigor and practical tool-building, which shaped university curricula at departments such as Department of Linguistics, University of Pennsylvania and language technology groups at Carnegie Mellon University. His legacy includes widely used corpora, pedagogical materials adopted by secondary and tertiary programs, and an expansive citation network connecting work at institutions like University of Illinois Urbana–Champaign, Peking University, and Seoul National University. He continues to influence corpus-based practice through software releases, datasets, and mentorship, leaving an enduring mark on lexicography and corpus linguistics worldwide.

Category:British linguists Category:Corpus linguistics