Skip to content
RAG Repo

VoxCeleb

VoxCeleb is a large speaker-recognition dataset built by the Visual Geometry Group at the University of Oxford. It gathers short speech segments from celebrity interviews on YouTube, spanning a wide range of accents, ages, and recording conditions, which makes it a realistic testbed for identifying and verifying who is speaking.

The dataset is a common benchmark for speaker verification (deciding whether two clips are the same person) and speaker identification. Note the split licence: the annotations Oxford provides are released under CC BY 4.0, but the audio itself comes from YouTube videos that remain under their original terms. That means the underlying content carries no blanket reuse right, so commercial use is restricted and you should check the source terms before building a product on it.

speechspeaker-recognitionyoutubevoiceresearch

Related sources