Skip to content
RAG Repo

GlotCC is aimed squarely at the long tail of world languages. Where most web-derived corpora concentrate on a handful of high-resource languages, GlotCC extends coverage to minority and low-resource ones, making it useful if you are building retrieval or generation for communities that are usually underserved by mainstream datasets.

The project releases both the corpus and the pipeline that produced it, so you can reproduce or extend the cleaning process yourself. The pipeline is licensed under CC0 (effectively public domain, no rights reserved), but the text itself comes from Common Crawl and remains subject to the original sites' terms, so the content licence differs from the pipeline licence and you should treat the underlying data accordingly.

multilinguallow-resourceminority-languagescommon-crawlcleaned

Related sources