Skip to content
RAG Repo

The People's Speech

The People's Speech is MLCommons' answer to a practical problem: most large speech datasets carry licences that make commercial use awkward. It gathers more than 30,000 hours of English audio from openly licensed sources and pairs it with transcripts, giving teams a big, legally usable corpus for training automatic speech recognition (ASR, turning spoken audio into text).

Licensing varies by subset: some parts are CC BY 4.0 and others are CC BY-SA 4.0. Both permit commercial use, but the ShareAlike (SA) portions require that derivative databases carry the same licence, so check which subset you are using before you build on it. The dataset is distributed through HuggingFace, which makes it straightforward to stream or download.

speechenglishasrcommercial-friendlymlcommons

Related sources