This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| VoxCeleb | |
|---|---|
| Name | VoxCeleb |
| Type | Speaker identification dataset |
| Creator | Oxford University Visual Geometry Group, DeepMind |
| Released | 2017 |
| Domain | Audio-visual speech, speaker recognition |
| Languages | Primarily English with multilingual samples |
| License | Research use |
VoxCeleb
VoxCeleb is a large-scale audio-visual speaker identification dataset developed to advance automatic speaker recognition and audio-visual association research. It provides thousands of utterances from hundreds to thousands of public figures collected from online videos, enabling work across speaker verification, diarization, and cross-modal retrieval. The corpus has been widely used by researchers at institutions such as Oxford University, DeepMind, Google Research, Facebook AI Research, Stanford University, and industrial labs including Amazon Research and Microsoft Research.
The corpus was introduced in a series of releases starting with VoxCeleb1 and expanded with VoxCeleb2, each increasing the number of identities and utterances and improving metadata. VoxCeleb1 drew from interviews, talk shows, and news clips featuring celebrities such as Barack Obama, Taylor Swift, Elon Musk, Oprah Winfrey, David Beckham, Beyoncé Knowles, Leonardo DiCaprio, and Angelina Jolie. VoxCeleb2 enlarged coverage to additional public figures including Jennifer Lawrence, Cristiano Ronaldo, Lionel Messi, Brad Pitt, Adele, Rihanna, Tom Cruise, Will Smith, Margaret Thatcher, Nelson Mandela, and many international personalities from Bollywood and K-pop circles. The dataset’s scale supports training deep convolutional and recurrent models developed in studies by groups like VGG, ResNet teams, and practitioners from MIT and Carnegie Mellon University.
Audio-visual samples were harvested from public video platforms featuring interviews, speeches, and panel discussions with identifiable speakers. Collections targeted high-profile events such as TED Conference talks, Academy Awards interviews, World Economic Forum sessions, United Nations addresses, and televised programs on networks like BBC, CNN, NBC, and Al Jazeera. Identity labels derive from automatic face detection and active speaker verification pipelines applied to footage of celebrities including Angela Merkel, Justin Bieber, Madonna, Kanye West, Stephen Hawking, J.K. Rowling, Vladimir Putin, Emmanuel Macron, Boris Johnson, and Jacinda Ardern. The corpus contains controlled splits for train, validation, and test across varied acoustic conditions and background noise drawn from urban settings, studio recordings, and live events featuring personalities such as Host of The Tonight Show guests.
Annotation combined automated pipelines and manual verification. Face tracking used models inspired by systems from FaceNet authors and research at Imperial College London; voice activity detection and speech segmentation benefitted from contributions by teams at CMU and Johns Hopkins University. Metadata includes identity tags, temporal boundaries, and confidence scores; utterances went through preprocessing like voice activity detection, bandpass filtering, and normalization that mirror practices in papers from Google DeepMind and Oxford VGG. For quality assurance, annotators cross-checked samples against source videos featuring figures such as Michelle Obama, Hillary Clinton, Pope Francis, Dalai Lama, Pelé, and Sachin Tendulkar to reduce label noise.
Standard evaluation protocols established speaker verification and identification tasks with trial lists and verification pairs. Researchers compared metric learning approaches, softmax-based classifiers, and angular margin losses popularized in work from Google Research and Facebook AI Research, reporting metrics like equal error rate (EER) and identification accuracy. Baselines include x-vector systems from teams at SRI International and end-to-end convolutional models employing architectures such as ResNet-34 and Inception. Leaderboards often reference experiments using additional corpora like LibriSpeech, TIMIT, and Switchboard to assess generalization across speakers including public figures like Ellen DeGeneres, Gordon Ramsay, Steven Spielberg, Martin Scorsese, Quentin Tarantino, and Hayao Miyazaki.
VoxCeleb catalyzed advances in speaker recognition, audio-visual diarization, and cross-modal retrieval used in systems by Apple, Google, Amazon, and academic projects at MIT Media Lab and Stanford Vision and Learning Lab. It supported research into robust embeddings used for forensic analysis, multimedia indexing for broadcasters like BBC and CNN, and assistive technologies for accessibility services implemented by teams at Microsoft and Facebook. The dataset has influenced competitions and benchmarks at venues like NeurIPS, ICASSP, INTERSPEECH, and CVPR where participants evaluated models on celebrity-centric test sets featuring personalities such as Arnold Schwarzenegger, Meryl Streep, Natalie Portman, Cate Blanchett, and Robert Downey Jr..
Because samples derive from public-domain videos of public figures, the dataset raises debates about consent, likeness rights, and surveillance implications discussed by ethicists at Harvard University, Yale University, and Stanford Law School. Critics cite potential misuse in deepfake generation or unauthorized identification of non-consenting speakers, prompting calls for governance frameworks like proposals discussed at European Parliament hearings and panels at ACM and IEEE. Mitigation strategies include restricted licenses, provenance metadata, and research-only access policies advocated by privacy researchers at Electronic Frontier Foundation and OpenAI.
VoxCeleb complements and inspired datasets such as LibriSpeech, Common Voice, TIMIT, Switchboard, AMI Corpus, MELD, and multimodal corpora from YouTube-8M. Successive efforts extended scale, diversity, and annotation richness in projects by Google Research, Facebook AI Research, and university consortia, producing specialized corpora for low-resource languages, speaker diarization challenges, and adversarial robustness evaluations featuring international and domain-specific speakers from sports, politics, and entertainment. Category:Speech recognition datasets