Skip to content
RAG Repo

MADLAD-400 (Multilingual And Document-level Large Audited Dataset) covers 419 languages, including many low-resource ones that most web corpora skip. It was built from Common Crawl and then manually audited, meaning people reviewed samples per language to catch mislabelled or low-quality text rather than relying on automated filters alone.

Two versions are available. The noisy version applies only language identification, so you get more data but with more junk. The cleaned version adds more extensive filtering for higher quality. Pick the noisy set if you want volume and plan to filter yourself, or the cleaned set if you want text that is closer to ready. Both are released under ODC-By 1.0, which permits commercial use with attribution.

multilingualpretrainingcommon-crawlgooglelow-resource

Related sources