LJ Speech
LJ Speech is a compact, clean corpus made for training and benchmarking text-to-speech (TTS, generating spoken audio from written text) systems. It consists of 13,100 short clips of one speaker reading passages from seven non-fiction books, with a transcript for each clip and durations between one and ten seconds.
Because it is a single speaker recorded consistently, it is easy to work with and has become the default starting point for TTS research and tutorials. The whole dataset is in the public domain, so you can use it for any purpose, including commercial work, with no restrictions or attribution required.
Related sources
LibriSpeech
A collection of 1,000 hours of read English speech, released under CC BY 4.0 and stored using the open-source FLAC audio encoder. Labels are aligned at the sentence level. Derived from LibriVox public domain audiobooks, it is the standard benchmark for English automatic speech recognition.
Mozilla Common Voice
A crowdsourced speech platform from Mozilla that releases datasets under the CC0 licence. Volunteers read sentences aloud and other community members validate each recording. As of release 19.0 it holds 32,584 hours of speech across 131 languages, making it one of the largest openly licensed voice datasets available.
OpenSLR
Open Speech and Language Resources, a hosting site for speech and language datasets, software, and models. It is the home of LibriSpeech and dozens of other language-specific speech corpora, making it a central catalogue for finding openly available voice data.
The People's Speech
A large, freely licensed English speech recognition dataset from MLCommons, assembled from openly licensed sources and totalling more than 30,000 hours. Built as a permissively licensed alternative for training commercial automatic speech recognition systems.