LLMpediaThe first transparent, open encyclopedia generated by LLMs

TED-LIUM

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: LibriSpeech Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

TED-LIUM
NameTED-LIUM
SubjectSpeech recognition dataset
Created2011
CreatorLaboratoire d'Informatique de l'Université du Maine, École Polytechnique Fédérale de Lausanne
LanguageEnglish
SizeAudio, transcriptions, alignments
LicenseMixed (see Licensing and Accessibility)

TED-LIUM is a publicly distributed corpus of English lecture-style audio derived from recorded TED Conference talks, assembled to support research in automatic speech recognition and spoken language processing. It aggregates audio, time-aligned transcriptions, speaker metadata, and linguistic resources suitable for training and evaluating acoustic models, language models, and end-to-end systems. The corpus has been used by researchers affiliated with organizations such as Google, Facebook, Microsoft Research, IBM Research, and academic groups at Massachusetts Institute of Technology, Stanford University, and Université Paris-Saclay.

Overview

The corpus compiles high-quality recordings from talks presented at TED Conference events by speakers including individuals connected to Harvard University, Princeton University, University of Cambridge, Stanford University, and University of Oxford. The project builds on practices established in corpora like Wall Street Journal (corpus), Switchboard (corpus), LibriSpeech, and CHiME (dataset), emphasizing reproducibility and standardized evaluation. Contributors to the dataset and its benchmarks have included researchers from INRIA, CNRS, Télécom ParisTech, École Normale Supérieure, and industry labs such as DeepMind and NVIDIA. The dataset is referenced in shared tasks and workshops hosted by venues such as INTERSPEECH, ICASSP, ACL (conference), and NeurIPS.

Dataset Versions and Contents

TED-LIUM has multiple releases that expand coverage and annotations: initial releases provided audio and orthographic transcriptions; later versions added forced alignments, speaker segmentation, and expanded talk counts. Major releases were prepared by teams at Laboratoire d'Informatique de l'Université du Maine and collaborators at LORIA (laboratory), and have been cited alongside corpora such as Common Voice and TEDx (series). Content per release typically includes waveform files, time-aligned phonetic alignments compatible with toolkits like Kaldi (software) and HTK (software), lexicons mapping to pronunciations used in CMU Pronouncing Dictionary, and metadata linking talks to presenters affiliated with institutions like Yale University, Columbia University, University of California, Berkeley, and University of Toronto.

Collection and Annotation Methodology

Audio was harvested from publicly available recordings of TED Conference talks, then processed to standard sampling rates and formats used by projects such as Mozilla Common Voice and Librispeech (dataset). Transcriptions were derived from published talk captions and human-corrected annotations following practices similar to those used in Fisher (corpus) and AMI Meeting Corpus. Forced alignment employed aligners compatible with Montreal Forced Aligner and toolkits like Kaldi (software), producing phone-level and word-level time stamps. Speaker labels and segmentation were generated using diarization pipelines inspired by research from SRI International, NIST, and groups involved in DIHARD (challenge). Quality control workflows referenced standards practiced at ELRA and LDC (Linguistic Data Consortium).

Licensing and Accessibility

Releases are distributed under terms that reflect TED content policies and third-party licensing considerations; access procedures resemble those for corpora managed by ELRA and LDC (Linguistic Data Consortium). Some components are available for direct download from hosting platforms coordinated by academic labs, while other parts require agreeing to redistribution constraints analogous to agreements for YouTube Corpus derivatives. Users in industry and academia often cite compliance checks comparable to those undertaken for datasets used by Google Research and Facebook AI Research. The dataset has been mirrored in repositories maintained by institutions such as INRIA and research groups at Université Grenoble Alpes.

Usage and Impact in Speech Recognition

TED-LIUM has been widely adopted as a benchmark in papers by teams at Carnegie Mellon University, University of Edinburgh, Johns Hopkins University, University of Illinois Urbana-Champaign, and companies like Amazon and Baidu. It has supported advances in hybrid HMM-DNN systems, end-to-end sequence-to-sequence models, and transformer-based architectures inspired by Google AI and the Vaswani et al. 2017 attention model. The corpus enabled comparisons across acoustic modeling approaches such as Deep Neural Networks, Convolutional Neural Networks, and Recurrent Neural Networks developed at labs including Facebook AI Research and DeepMind, and influenced language modeling work leveraging corpora like Wikipedia and Common Crawl.

Benchmark Results and Evaluation Protocols

Standard evaluation protocols define training, development, and test splits; scoring metrics include word error rate (WER) and time-aligned scoring used in evaluation campaigns by NIST and benchmarks published at conferences like ICASSP and INTERSPEECH. Reported results in the literature span baseline GMM-HMM systems to state-of-the-art end-to-end models evaluated by researchers at Kyoto University, University of Montreal, ETH Zurich, and corporate research teams at Microsoft Research and Apple Machine Learning Research. Reproducible recipes for toolkits such as Kaldi (software) and libraries like PyTorch and TensorFlow accompany many papers to standardize comparisons.

Associated resources include pronunciation lexicons, language modeling text extracted from talk transcripts linked to institutions like MIT Press authors, and forced-alignment outputs compatible with toolchains from Montreal Tools and Kaldi (software). Extensions and derivative datasets connect TED-LIUM to projects like LibriSpeech, Common Voice, AMI Meeting Corpus, and domain-specific corpora used by DARPA programs. Ongoing community efforts integrate TED-LIUM with multilingual initiatives involving ELRA, LDC (Linguistic Data Consortium), and academic consortia at University of Cambridge and University of Oxford.

Category:Speech recognition datasets