Skip to content
RAG Repo

Dolma

Dolma pulls together several sources into one documented corpus: web text from Common Crawl, scientific papers from Semantic Scholar, code from GitHub, public-domain books, Reddit conversations, and Wikipedia. It spans more than 4 billion documents totalling 3T tokens, so the mix inside it is genuinely varied, hence its `mixed` readiness rating.

What sets Dolma apart is transparency. The Allen Institute for AI (Allen AI) published the full toolkit used to build it, including the filtering, deduplication, and mixing steps, so you can see and reproduce how every part was assembled. It was created as the training data for the OLMo family of fully open models.

It is hosted on HuggingFace and downloadable in shards. Note the licensing is split: ODC-By 1.0 alongside the AI2 ImpACT License, so check the terms for the subset you plan to use.

web-crawlpretrainingmulti-sourceopen-corpusallen-ai

Related sources