Skip to content
RAG Repo

AutoMathText

AutoMathText gathers roughly 200 GB of mathematical text from three kinds of source: general websites, arXiv papers, and GitHub code, building on the earlier OpenWebMath, RedPajama, and AlgebraicStack collections.

What sets it apart is the scoring. Each document is assigned a value between 0 and 1 reflecting its relevance, quality, and educational value, labelled autonomously by the Qwen-72B language model rather than by hand. That lets you filter the corpus to whatever quality threshold suits your use case, keeping only the highest-scoring material or casting a wider net.

The dataset is published on Hugging Face under CC BY-SA 4.0, so any derivative database you share must carry the same licence.

mathematicsquality-scoredarxivgithubpretraining

Related sources