LLMpediaThe first transparent, open encyclopedia generated by LLMs

AMI Meeting Corpus

Note: This article was automatically generated by a large language model (LLM) from purely parametric knowledge (no retrieval). It may contain inaccuracies or hallucinations. This encyclopedia is part of a research project currently under review.
Article Genealogy
Parent: LibriSpeech Hop 5 terminal

This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.

AMI Meeting Corpus
NameAMI Meeting Corpus
Other namesAMI Corpus
SubjectMultimodal meeting recordings
CreatorsAMI Project
Released2006
FormatAudio, video, transcripts, annotations
AccessResearch use

AMI Meeting Corpus

The AMI Meeting Corpus is a multimodal dataset of recorded meetings used for research in speech recognition, natural language processing, computer vision, human–computer interaction, and machine learning. Developed by the AMI Project and affiliated institutions, it contains synchronized audio and video streams alongside manual transcripts and rich annotations for analysis of dialogue, speaker diarization, gesture, and turn-taking. The corpus has been widely used in evaluations, workshops, and challenges organized by groups in Europe and North America.

Overview

The corpus was created through collaborations involving the European Union research frameworks, the Idiap Research Institute, King's College London, University of Sheffield, Rothamsted Research, and industrial partners including Microsoft Research and IBM Research. Recordings emulate real-world meeting settings inspired by scenarios from companies such as Siemens, Nokia, Philips, BT Group, and Hewlett-Packard. Collections were influenced by prior corpora like TIMIT, Switchboard Corpus, CHiME Challenge datasets, and later compared with datasets such as LibriSpeech and CHILDES in multimodal studies.

Corpus Composition

The dataset comprises approximately 100 hours of meetings with audio captured by close-talk and far-field microphones, video from multiple cameras, and data from microphones arrays and wearable devices. Sessions include four-person scenario meetings covering design reviews, brainstorming, and decision-making; participant roles reflect typical organizational settings found at Cambridge University, Oxford University, Imperial College London, and research labs such as Bell Labs and SRI International. Annotations encompass manual orthographic transcripts, time-aligned word-level timestamps, dialogue act tags, topic segmentation, and gesture labels; these annotation schemes align with standards used by groups at NIST, ACL, IEEE, and ISCA.

Data Collection and Annotation

Collection protocols followed ethical review procedures common at institutions like ETH Zurich and University of Edinburgh, with informed consent and controlled environments arranged in venues similar to conference rooms at Googleplex, Apple Park, and university facilities. Annotators employed tools and schemas developed in collaboration with projects at Stanford University, MIT, and Carnegie Mellon University to mark disfluencies, overlaps, and speaker turns. Transcription guidelines drew on conventions used by LDC and ELRA; annotations were further cross-validated by teams from Université Paris-Sud and KU Leuven.

Access and Licensing

The corpus is distributed for research purposes under licensing terms negotiated with entities like British Library and data archiving agencies such as ELRA and LDC; access typically requires a data agreement and adherence to privacy provisions modeled after policies at EU GDPR-influenced institutions. Requests for access have been coordinated through university repositories and research consortia including CLARIN and DARIAH, and usage has been tracked in publications from ICML, NeurIPS, ICASSP, INTERSPEECH, and EMNLP.

Applications and Impact

Researchers have applied the data to advance automatic speech recognition models, improve speaker identification, and develop multimodal fusion algorithms combining audio, video, and gesture signals. Work using the corpus influenced systems from Amazon, Facebook AI Research, Google DeepMind, and academic labs at University of Toronto and University of California, Berkeley. The dataset underpinned shared tasks at venues like SIGdial, Interspeech, ICASSP, and competitions supported by DARPA and shaped benchmarks in studies by teams at Microsoft, Siemens Research, Xerox PARC, and Adobe Research.

Evaluation and Benchmarks

AMI has been included in benchmark suites for diarization, speech recognition, and meeting summarization, with performance comparisons reported in proceedings of ACL, EMNLP, NAACL, SLT Workshop, and workshops at CVPR and ECCV. Baseline systems were developed by consortia involving Cambridge University Engineering Department, IDIAP, TNO, and Baidu Research, and evaluated against metrics standardized by NIST and scoring methods used at MIREX and ROUGE competitions. Results informed model improvements in neural architectures by groups at DeepMind, OpenAI, FAIR, and universities such as Peking University.

Limitations and Criticisms

Critiques have addressed representativeness, noting that scripted scenarios and participant demographics differ from meetings in corporations like Goldman Sachs or governmental bodies such as United Nations assemblies. Concerns about privacy, annotation consistency, and cross-cultural generalization were raised by scholars at Princeton University, Yale University, Columbia University, and University of Michigan. The dataset's recording conditions and microphone configurations limit direct applicability to in-the-wild settings like crowded venues at La Défense or open-plan offices modeled after WeWork locations. Subsequent corpora from initiatives at Facebook, Google, and Amazon have sought to address some of these gaps.

Category:Speech corpora