This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| Open Speech and Language Resources | |
|---|---|
| Name | Open Speech and Language Resources |
| Focus | Speech recognition, natural language processing, corpora, datasets |
| Region | Global |
| Established | Various initiatives since 1990s |
Open Speech and Language Resources are publicly accessible datasets, corpora, tools, and platforms used for research and development in speech recognition and natural language processing. They underpin experimental work across industry and academia, supporting benchmarks, multilingual research, and reproducibility in fields influenced by initiatives associated with Alan Turing, Claude Shannon, Noam Chomsky, John McCarthy, and Geoffrey Hinton. Major projects intersect with efforts by Massachusetts Institute of Technology, Stanford University, Carnegie Mellon University, University of Cambridge, and University of Oxford.
Open Speech and Language Resources encompass diverse materials created by institutions such as Google, Microsoft, Facebook, Amazon, IBM, Apple Inc., Baidu, Tencent, Huawei, NVIDIA, OpenAI, DeepMind, Mozilla Foundation, Mozilla, and research groups at Harvard University, Princeton University, Yale University, Columbia University, University of California, Berkeley, and University of Toronto. They relate to standards and evaluation campaigns run by Linguistic Data Consortium, European Language Resources Association, National Institute of Standards and Technology, DARPA, European Commission, Tokyo Institute of Technology, and Korea Advanced Institute of Science and Technology. Historical precedents include datasets and corpora connected to projects at Bell Labs, Bell Telephone Company, AT&T, Xerox PARC, SRI International, and initiatives supported by National Science Foundation, Wellcome Trust, and Horizon 2020.
Common categories include speech corpora, text corpora, lexicons, pronunciation dictionaries, parallel corpora, annotated datasets, dialogue corpora, conversational agents, and toolkits developed by organizations like Allen Institute for AI, European Language Grid, Common Voice, Wikimedia Foundation, Project Gutenberg, Internet Archive, British Library, Library of Congress, National Diet Library (Japan), and Deutsche Nationalbibliothek. Resources are often created in collaboration with research centers at McGill University, ETH Zurich, Imperial College London, University of Edinburgh, Georgia Institute of Technology, Peking University, Tsinghua University, Indian Institute of Technology Bombay, University of Delhi, University of Melbourne, Australian National University, and University of Sydney.
Prominent open datasets and corpora originate from projects and institutions such as Linguistic Data Consortium releases, European Language Resources Association catalogs, Common Voice by Mozilla Foundation, LibriSpeech (derived from Project Gutenberg and LibriVox), TED-LIUM associated with TED Conferences, and corpora used in Wall Street Journal-based research at AT&T Bell Labs. Other notable mentions include datasets influenced by work at Google Research, Microsoft Research, Facebook AI Research, DARPA, NIST, CHiME Challenge, ICASSP proceedings, and benchmarks used by teams at DeepMind, OpenAI, NVIDIA, Adobe Inc., Samsung Electronics, and ARM Holdings.
Licensing regimes for open resources draw from frameworks used by Creative Commons, Open Data Commons, GNU Project, Free Software Foundation, MIT License, Apache License, European Union Agency for Fundamental Rights, United Nations Educational, Scientific and Cultural Organization, and regulatory bodies including European Parliament directives and rulings by courts in United States, European Union, United Kingdom, Canada, Australia, and Japan. Ethical considerations reference guidelines from Association for Computational Linguistics, IEEE Standards Association, ACM, Human Rights Watch, Amnesty International, Electronic Frontier Foundation, and policy discussions involving United Nations, World Health Organization, and Organisation for Economic Co-operation and Development.
Key tools and frameworks for processing open resources are developed by TensorFlow, PyTorch, Kaldi, ESPnet, HTK (Hidden Markov Model Toolkit), Julius (speech recognition engine), OpenFST, KenLM, Fairseq, Hugging Face, SpaCy, NLTK, Gensim, and platforms maintained by GitHub, GitLab, Bitbucket, Google Cloud Platform, Amazon Web Services, Microsoft Azure, and IBM Cloud. Integrations often reference standards from World Wide Web Consortium, Unicode Consortium, and multilingual initiatives tied to European Language Resources Association and UNESCO.
Community initiatives and governance structures are spearheaded by consortia and projects including Mozilla Foundation's initiatives, Wikimedia Foundation collaborations, Open Knowledge Foundation, Data Carpentry, Software Carpentry, Carnegie Mellon University centers, Stanford Human-Centered AI Institute, Alan Turing Institute, Center for Data Innovation, and European projects funded by Horizon 2020 and managed through partnerships with European Commission. Oversight and standards involve participation from Association for Computational Linguistics, IEEE, ACM, National Institutes of Health, Wellcome Trust, Bill & Melinda Gates Foundation, and philanthropic actors like Open Philanthropy Project.
Key challenges involve data diversity, bias, and representation highlighted in work from ProPublica, New York Times, The Guardian, and research by scholars at MIT Media Lab, Oxford Internet Institute, Berkman Klein Center for Internet & Society, and AI Now Institute. Future directions point to multilingual expansion championed by United Nations, language preservation efforts allied with Smithsonian Institution, Endangered Languages Project, collaborations with governments in India, China, Brazil, Nigeria, and research funded by agencies such as DARPA, European Research Council, National Science Foundation, and Japan Science and Technology Agency.
Category:Speech recognition Category:Natural language processing Category:Open data