OpenSLR
OpenSLR is a long-running hosting site that gathers speech and language resources in one place: datasets, software, and trained models. Each resource gets a numbered slot (LibriSpeech is number 12, for example), and the catalogue spans dozens of languages, from widely spoken ones to smaller corpora that are hard to find elsewhere.
It is best thought of as a directory and download host rather than a single dataset. Licensing varies from one resource to the next, so check the terms on each individual entry before you use it: some are permissive, others carry attribution or ShareAlike conditions. If you are hunting for speech data in a specific language, this is one of the first places worth checking.
Related sources
LibriSpeech
A collection of 1,000 hours of read English speech, released under CC BY 4.0 and stored using the open-source FLAC audio encoder. Labels are aligned at the sentence level. Derived from LibriVox public domain audiobooks, it is the standard benchmark for English automatic speech recognition.
LJ Speech
An open dataset of 13,100 short audio clips of a single speaker reading passages from seven non-fiction books. Every clip is transcribed and clips run from 1 to 10 seconds. Released into the public domain, it is the standard single-speaker text-to-speech benchmark.
Mozilla Common Voice
A crowdsourced speech platform from Mozilla that releases datasets under the CC0 licence. Volunteers read sentences aloud and other community members validate each recording. As of release 19.0 it holds 32,584 hours of speech across 131 languages, making it one of the largest openly licensed voice datasets available.
The People's Speech
A large, freely licensed English speech recognition dataset from MLCommons, assembled from openly licensed sources and totalling more than 30,000 hours. Built as a permissively licensed alternative for training commercial automatic speech recognition systems.