<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>RAG Repo - What&apos;s new</title><description>Recently added and reviewed data sources in the RAG Repo directory.</description><link>https://rag-repo.org/</link><item><title>CASE Network (1EdTech)</title><link>https://rag-repo.org/source/case-network/</link><guid isPermaLink="true">https://rag-repo.org/source/case-network/</guid><description>A public registry of machine-readable learning-standard frameworks from all 50 US states and other issuing agencies, run by 1EdTech in the CASE JSON format. The digitally referenceable spine of what to teach, at which level and in what order, rather than the teaching content itself.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>CK-12 Foundation</title><link>https://rag-repo.org/source/ck-12/</link><guid isPermaLink="true">https://rag-repo.org/source/ck-12/</guid><description>Free K-12 library of customisable FlexBook textbooks, strong in maths and science, with adaptive practice, interactive simulations and study guides, organised by grade level and aligned to state standards. Non-commercial licence.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Common Core State Standards</title><link>https://rag-repo.org/source/common-core/</link><guid isPermaLink="true">https://rag-repo.org/source/common-core/</guid><description>The US Common Core State Standards for English language arts and mathematics, defined grade by grade from kindergarten to grade twelve. A widely adopted curriculum framework, available in machine-readable form through ASN and CASE.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Khan Academy</title><link>https://rag-repo.org/source/khan-academy/</link><guid isPermaLink="true">https://rag-repo.org/source/khan-academy/</guid><description>A free nonprofit learning platform organised as a subject-to-skill knowledge tree across maths, science, economics and the humanities, from primary through early college. Videos, articles and mastery-based exercises, but non-commercial licensing and no public data API.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Next Generation Science Standards</title><link>https://rag-repo.org/source/ngss/</link><guid isPermaLink="true">https://rag-repo.org/source/ngss/</guid><description>US K-12 science standards developed by a consortium of states and released in 2013. A curriculum framework of performance expectations arranged by grade band and disciplinary core idea, free to use with attribution and browsable online or downloadable as PDFs.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Saylor Academy</title><link>https://rag-repo.org/source/saylor-academy/</link><guid isPermaLink="true">https://rag-repo.org/source/saylor-academy/</guid><description>A nonprofit offering more than 300 free, self-paced college-level online courses across many subjects, with curated readings, framing text, and assessments. Course outlines are CC BY licensed; embedded third-party materials keep their own varying licences.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Siyavula</title><link>https://rag-repo.org/source/siyavula/</link><guid isPermaLink="true">https://rag-repo.org/source/siyavula/</guid><description>Free, CAPS-aligned Everything Maths and Everything Science textbooks from Siyavula, covering South African school grades. Read online or download as PDF and EPUB, openly licensed for copying and redistribution.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>UK National Curriculum (England)</title><link>https://rag-repo.org/source/uk-national-curriculum/</link><guid isPermaLink="true">https://rag-repo.org/source/uk-national-curriculum/</guid><description>The statutory programmes of study and attainment targets for state schools in England, organised by key stage and subject and published by the Department for Education. The curriculum spine for England, freely reusable under the Open Government Licence.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Wikiversity</title><link>https://rag-repo.org/source/wikiversity/</link><guid isPermaLink="true">https://rag-repo.org/source/wikiversity/</guid><description>Free, collaboratively written learning materials from the Wikimedia Foundation: lessons, courses and tutorials organised by subject and educational level, in many languages. Available as bulk dumps and via API, like the other Wikimedia projects.</description><pubDate>Tue, 15 Sep 2026 00:00:00 GMT</pubDate></item><item><title>AlphaFold Protein Structure Database</title><link>https://rag-repo.org/source/alphafold-protein-structure-database/</link><guid isPermaLink="true">https://rag-repo.org/source/alphafold-protein-structure-database/</guid><description>EMBL-EBI and Google DeepMind&apos;s open database of AI-predicted protein structures, covering over 200 million sequences across almost all of UniProt. Each model ships with per-residue confidence and predicted aligned error scores, and is reachable by website, API, FTP and Google Cloud bulk access. Released under CC BY 4.0 and designed to pair with UniProt and the experimental structures in the PDB.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Berkeley Function-Calling Leaderboard (BFCL)</title><link>https://rag-repo.org/source/bfcl/</link><guid isPermaLink="true">https://rag-repo.org/source/bfcl/</guid><description>The de facto benchmark for how well language models call functions, APIs and tools. Built by UC Berkeley&apos;s Gorilla project, it spans Python, Java, JavaScript and REST with simple, parallel, irrelevance-detection, multi-turn and agentic cases. Apache 2.0 and freely available.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>BIGPATENT</title><link>https://rag-repo.org/source/bigpatent/</link><guid isPermaLink="true">https://rag-repo.org/source/bigpatent/</guid><description>A corpus of 1.3 million US utility patents filed between 1971 and 2018, each paired with its human-written abstract as a gold-standard summary and organised by Cooperative Patent Classification code. A large, clean patent text corpus built for abstractive summarisation and other patent NLP work.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Essential-Web v1.0 (v1.0)</title><link>https://rag-repo.org/source/essential-web/</link><guid isPermaLink="true">https://rag-repo.org/source/essential-web/</guid><description>A 24-trillion-token web corpus (about 23.6 billion documents) from Essential AI where every document carries a twelve-category taxonomy covering topic, format, complexity and quality. The labels let you carve out domain or quality subsets with simple filters, without training your own classifiers.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>FinePDFs</title><link>https://rag-repo.org/source/finepdfs/</link><guid isPermaLink="true">https://rag-repo.org/source/finepdfs/</guid><description>HuggingFace&apos;s corpus built entirely from PDFs: roughly 3 trillion tokens across 475 million documents in 1,733 languages, drawn from Common Crawl and the open web. It captures dense long-form content (reports, manuals, papers) that HTML web corpora miss, and is the PDF counterpart to FineWeb.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>FineWeb-Edu</title><link>https://rag-repo.org/source/fineweb-edu/</link><guid isPermaLink="true">https://rag-repo.org/source/fineweb-edu/</guid><description>The educational-quality subset of FineWeb, about 1.3 trillion English tokens kept by an educational-quality classifier, with a larger 5.4T token variant at a looser threshold. Built by HuggingFace as a denser, more teachable base for pretraining and broad-coverage RAG.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>FreshStack</title><link>https://rag-repo.org/source/freshstack/</link><guid isPermaLink="true">https://rag-repo.org/source/freshstack/</guid><description>A 2025 framework and benchmark for retrieval over fast-moving technical documentation. It pairs real Stack Overflow questions and answers with chunked GitHub code and docs across five niche domains, giving code and docs RAG a hard, contamination-resistant test.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Glaive Function Calling v2 (v2)</title><link>https://rag-repo.org/source/glaive-function-calling/</link><guid isPermaLink="true">https://rag-repo.org/source/glaive-function-calling/</guid><description>A widely used open dataset of about 113,000 synthetic multi-turn chat conversations that include function calls and their results, made by Glaive AI. One of the most downloaded open datasets for fine-tuning models to call tools, released under Apache 2.0.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Google-Microsoft-OSM Open Buildings (VIDA)</title><link>https://rag-repo.org/source/google-microsoft-open-buildings/</link><guid isPermaLink="true">https://rag-repo.org/source/google-microsoft-open-buildings/</guid><description>A conflated global building-footprint layer that merges Google Open Buildings V3, Microsoft Global ML Footprints, and OpenStreetMap into roughly 2.7 billion footprints, each labelled by its source. Built by VIDA and hosted on Source Cooperative in cloud-native formats, it offers more complete coverage than any single provider on its own.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>LegalBench-RAG</title><link>https://rag-repo.org/source/legalbench-rag/</link><guid isPermaLink="true">https://rag-repo.org/source/legalbench-rag/</guid><description>The first open benchmark for the retrieval step of legal RAG. It offers 6,858 human-annotated query-answer pairs over a corpus of roughly 79 million characters, with character-level ground-truth spans across contracts and privacy policies, assembled from CUAD, MAUD, ContractNLI and PrivacyQA.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Nemotron Post-Training Dataset v2 (v2)</title><link>https://rag-repo.org/source/nemotron-post-training-v2/</link><guid isPermaLink="true">https://rag-repo.org/source/nemotron-post-training-v2/</guid><description>NVIDIA&apos;s 2025 post-training set of prompts and synthetic responses for supervised fine-tuning and reinforcement learning, covering mathematics, code, STEM, reasoning, and instruction following, with the instruction data expanded into five additional languages.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Open Materials 2024 (OMat24) (2024)</title><link>https://rag-repo.org/source/omat24/</link><guid isPermaLink="true">https://rag-repo.org/source/omat24/</guid><description>Meta FAIR&apos;s open dataset of more than 110 million DFT calculations of inorganic materials, sampled for structural and compositional diversity. Models trained on it top the MatBench-Discovery leaderboard for stability and formation-energy prediction.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>OpenThoughts3-1.2M</title><link>https://rag-repo.org/source/openthoughts3/</link><guid isPermaLink="true">https://rag-repo.org/source/openthoughts3/</guid><description>An open reasoning dataset of 1.2 million rows (roughly 850k maths, 250k code and 100k science questions) each paired with a chain-of-thought answer. Released by the Open Thoughts team in June 2025, it is the training set behind OpenThinker3 and is licensed Apache 2.0.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>RCSB Protein Data Bank</title><link>https://rag-repo.org/source/rcsb-protein-data-bank/</link><guid isPermaLink="true">https://rag-repo.org/source/rcsb-protein-data-bank/</guid><description>The canonical open archive of experimentally determined 3D biomolecular structures, over 250,000 of them, served alongside more than a million computed structure models. It is the ground truth of structural biology, released into the public domain under CC0.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>RTEB (Retrieval Embedding Benchmark)</title><link>https://rag-repo.org/source/rteb/</link><guid isPermaLink="true">https://rag-repo.org/source/rteb/</guid><description>The MTEB team&apos;s retrieval-focused embedding benchmark, launched in beta in October 2025 as a new retrieval section of the MTEB leaderboard. It spans 20 languages and enterprise domains such as law, healthcare, finance and code, and deliberately mixes open datasets with held-out private ones to measure genuine generalisation rather than training-set memorisation.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Structured Wikipedia</title><link>https://rag-repo.org/source/structured-wikipedia/</link><guid isPermaLink="true">https://rag-repo.org/source/structured-wikipedia/</guid><description>Wikipedia rendered as pre-parsed, machine-readable JSON: abstracts, short descriptions, infoboxes, sections, parsed tables and references, with links to Wikidata entities. A beta from Wikimedia Enterprise covering nine languages, also mirrored on Hugging Face. The section-segmented shape a RAG pipeline actually wants.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>SYNTH</title><link>https://rag-repo.org/source/synth-pleias/</link><guid isPermaLink="true">https://rag-repo.org/source/synth-pleias/</guid><description>A fully open synthetic corpus of amplified multilingual encyclopaedic text with built-in reasoning traces and exercises covering RAG, information extraction and QA. Released by PleIAs with the AI Alliance, it targets training and evaluating small, grounded, citeable reasoning models.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>T²-RAGBench</title><link>https://rag-repo.org/source/t2-ragbench/</link><guid isPermaLink="true">https://rag-repo.org/source/t2-ragbench/</guid><description>A 2025 benchmark of 23,088 context-independent question, context and answer triples over more than 7,300 financial documents that mix text and tables. Each question maps to exactly one ground-truth document, making it purpose-built for numerical, table-aware financial RAG.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>The Stack v2 (v2)</title><link>https://rag-repo.org/source/the-stack-v2/</link><guid isPermaLink="true">https://rag-repo.org/source/the-stack-v2/</guid><description>BigCode&apos;s large-scale source code corpus and the pretraining set behind StarCoder2, built with Software Heritage and spanning more than 600 programming languages. The full version runs to about 67.5TB, with records pointing to code held in Software Heritage&apos;s S3 rather than embedding it directly.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>ToolBench</title><link>https://rag-repo.org/source/toolbench/</link><guid isPermaLink="true">https://rag-repo.org/source/toolbench/</guid><description>An open instruction-tuning dataset for teaching general tool use to language models, built for the ToolLLM project over 16,464 real-world REST APIs from RapidAPI across 49 categories, with single-tool and multi-tool, multi-step solution paths.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item><item><title>Toucan-1.5M</title><link>https://rag-repo.org/source/toucan-agentic/</link><guid isPermaLink="true">https://rag-repo.org/source/toucan-agentic/</guid><description>The largest open tool-agentic dataset: over 1.5 million trajectories synthesised from 495 real-world MCP servers spanning 2,000 plus tools, with multi-turn, sequential and parallel tool calls backed by real executions and error handling. Built by Agent-Ark and released under Apache 2.0, it is premier open data for training retrieval-and-tool (MCP) agents.</description><pubDate>Thu, 10 Sep 2026 00:00:00 GMT</pubDate></item></channel></rss>