AI4Bharat (IndicCorp)
A collection of corpora, models, and benchmarks for Indian languages, produced by IIT Madras. It covers major Indic languages with monolingual corpora, parallel translation data, and evaluation sets.
Large-scale text corpora covering languages beyond English, including low-resource and regional languages. Essential for building RAG systems for non-English or multilingual audiences.
9sources
A collection of corpora, models, and benchmarks for Indian languages, produced by IIT Madras. It covers major Indic languages with monolingual corpora, parallel translation data, and evaluation sets.
A high-quality multilingual corpus derived from Common Crawl. It splits every document into separate paragraphs, so it is effectively a paragraph-level rather than document-level corpus, which matters for your chunking strategy.
A multilingual text dataset of 6.3 T tokens across 167 languages, built for training large language models. It combines cleaned versions of mC4 and OSCAR, two large web corpora, after a multi-stage filtering pipeline.
An open, broad-coverage corpus and processing pipeline built from Common Crawl, targeting minority and low-resource languages that larger corpora tend to miss. It fills gaps left by datasets focused on higher-resource languages.
An EU-funded project building open language resources from web crawls and the Internet Archive. Its MonoHPLT dataset contains monolingual text in 75 languages, released as openly as licensing allows.
A manually audited, general-domain monolingual dataset of 3 T tokens spanning 419 languages, built from Common Crawl by Google DeepMind and Google Research. It comes in a noisy version and a more heavily filtered clean version.
A grassroots research community building natural language processing resources for African languages. It produces datasets, benchmarks, and models across dozens of African languages that are otherwise poorly served.
A large corpus of translated movie and TV subtitles. The latest version covers 60 languages with 2.6 billion sentences in total. It is valuable for conversational and colloquial language, and for parallel translation data.
The Responsible Open-Science Open-Collaboration Text Sources corpus, a 1.6 TB dataset spanning 46 natural languages and 13 programming languages, released by BigScience. Built from community-selected sources plus OSCAR, it trained BLOOM.