Skip to content
RAG Repo

OpenWebMath

OpenWebMath sets out to capture the bulk of the high-quality mathematical text on the open web. The team started from more than 200 billion HTML files in the Common Crawl archive and applied maths-aware extraction and filtering to keep the pages that carry real mathematical content, preserving equations and notation that generic web extractors tend to mangle.

The result is 6.3 million documents worth roughly 14.7 billion tokens. While the majority of the content is mathematics, a large share covers related technical subjects, including physics, computer science, and statistics, so it works well as a broad reasoning corpus rather than a narrow one.

The dataset is published on Hugging Face and is a common ingredient in larger maths corpora such as Proof-Pile-2 and AutoMathText.

mathematicsweb-crawlcommon-crawlpretrainingphysicsenglish

Related sources