Browse all datasets
Every source in one place. Search by name, description, or tag, and filter by category, access type, and how ready the data is for retrieval.
More filters: RAG readiness
Search results
238 sources
- Multilingual & regional corporaOpen
AI4Bharat (IndicCorp)
A collection of corpora, models, and benchmarks for Indian languages, produced by IIT Madras. It covers major Indic languages with monolingual corpora, parallel translation data, and evaluation sets.
multilingualindic-languagesindia - Biomedical & healthOpen
AlphaFold Protein Structure Database
EMBL-EBI and Google DeepMind's open database of AI-predicted protein structures, covering over 200 million sequences across almost all of UniProt. Each model ships with per-residue confidence and predicted aligned error scores, and is reachable by website, API, FTP and Google Cloud bulk access. Released under CC BY 4.0 and designed to pair with UniProt and the experimental structures in the PDB.
proteinsstructure-predictionbioinformatics - Mathematics & reasoningOpen
AMPS
A dataset of informal mathematics introduced alongside the MATH benchmark. It includes more than 100,000 Khan Academy problems with step-by-step solutions in LaTeX and over 5 million problems generated with Mathematica scripts, totalling around 23 GB.
mathematicsproblem-solvinglatex - Academic & scientific literatureOpen
arXiv
An open-access preprint server for physics, mathematics, computer science, quantitative biology, statistics, and more, holding over 2.5 million papers. You can pull the full archive in bulk from Amazon S3, or harvest the metadata through OAI-PMH, a standard protocol for sharing records between repositories.
preprintsphysicsmathematics - Legal & regulatoryOpen
AustLII
A free resource of medium-neutral case law and unreported judgments for all Australian jurisdictions, covering the whole country since 1995.
case-lawaustralialegislation - Mathematics & reasoningOpen
AutoMathText
Around 200 GB of mathematical text compiled from websites, arXiv, and GitHub, drawing on OpenWebMath, RedPajama, and AlgebraicStack. Every piece of content carries a score from 0 to 1 for relevance, quality, and educational value, labelled automatically by the Qwen-72B model.
mathematicsquality-scoredarxiv - Mathematics & reasoningOpen
Awesome AI Math Datasets
A community-curated list of open-source mathematics datasets for training and evaluating maths-capable language models. A useful index for finding newer additions in this space.
mathematicscurated-listmeta-resource - Curated lists & meta-resourcesOpen
Awesome Legal Data
A community-maintained list of legal datasets, tools, and resources for legal text processing across jurisdictions, including court records, statutes, contracts, and legal NLP benchmarks. A useful map for anyone building a legal RAG system.
awesome-listlegaldirectory - Legal & regulatoryOpen
Awesome LegalTech
A curated list of legal technology resources: open-source platforms, AI models, companies, datasets, and tools spanning the global legal ecosystem. Useful for tracking new entrants in a fast-moving space.
meta-resourcecurated-listlegaltech - Curated lists & meta-resourcesOpen
Awesome Public Datasets
A community-curated list of high-quality open datasets on GitHub, organised by topic: agriculture, biology, climate, economics, education, finance, government, healthcare, and more. A good starting point when you need RAG-ready data for a specific domain and do not yet know where to look.
awesome-listdirectoryopen-data - Data platforms & marketplacesCommercial
AWS Data Exchange
A marketplace for finding, subscribing to, and using third-party data inside the AWS cloud. It carries both free open datasets and paid commercial data products, so you can pull licensed data straight into your AWS workflows without setting up separate transfers.
data-platformmarketplacecommercial - Data platforms & marketplacesOpen
AWS Open Data Registry
A registry of high-value datasets hosted on AWS and made publicly available, covering genomics, geospatial data, climate, satellite imagery, and more. Over 300 PB of data in total, free to access: you pay only for the compute you use to process it.
data-platformcloudgeospatial - Legal & regulatoryOpen
BAILII (British and Irish Legal Information Institute)
Free access to British and Irish primary legal materials, covering UK and Ireland case law and legislation. Alongside the National Archives, one of the main free sources for reading UK judgments.
case-lawukireland - Retrieval benchmarks & evaluationOpen
BEIR
A collection of 18 information retrieval datasets spanning ad-hoc web search, question answering, fact verification, and duplicate question retrieval. Built by aggregating existing datasets, some originally created for other tasks and converted to retrieval format. Tests how well retrieval models generalise to unseen domains without fine-tuning, which matters for real-world RAG.
retrievalbenchmarkinformation-retrieval - Agentic & Tool-UseOpen
Berkeley Function-Calling Leaderboard (BFCL)
The de facto benchmark for how well language models call functions, APIs and tools. Built by UC Berkeley's Gorilla project, it spans Python, Java, JavaScript and REST with simple, parallel, irrelevance-detection, multi-turn and agentic cases. Apache 2.0 and freely available.
function-callingtool-useagentic - Patents & intellectual propertyOpen
BIGPATENT
A corpus of 1.3 million US utility patents filed between 1971 and 2018, each paired with its human-written abstract as a gold-standard summary and organised by Cooperative Patent Classification code. A large, clean patent text corpus built for abstractive summarisation and other patent NLP work.
patentsususpto - Chemistry, materials & life sciencesOpen
Bio2RDF
An open-source project that pulls together a diverse set of life-sciences datasets from many providers into a single linked-data graph, with a SPARQL endpoint for querying across them. The full collection is about 11 billion triples across 35 datasets, including DrugBank, PubMed, and MeSH.
life-sciencesknowledge-graphsparql - Retrieval benchmarks & evaluationOpen
BRIGHT
A reasoning-intensive retrieval benchmark. It requires genuine reasoning to connect queries with relevant documents rather than keyword or semantic overlap, revealing model weaknesses that BEIR misses.
retrievalbenchmarkreasoning - Web corporaOpen
C4 (Colossal Clean Crawled Corpus)
A 750 GB English corpus derived from Common Crawl using heuristic cleaning to keep natural language and drop gibberish, boilerplate, and placeholder text. Built to train Google's T5 and later used for MPT-7B and others.
web-crawlenglishpretraining - Legal & regulatoryLimited
Cambridge Law Corpus
A research dataset of more than 250,000 UK court cases, mostly from the 21st century but with some reaching back to the 16th century. Built by the University of Cambridge for legal natural language processing work, it is available under restricted access for research use.
ukcase-lawlegal - Legal & regulatoryLimited
CanLII
A national legal information institute covering Canada's federal and provincial jurisdictions, hosting over 300 databases of legislation and case law. Bilingual, with thousands of commentaries on Canadian court decisions.
case-lawcanadalegislation - Education & open learningOpen
CASE Network (1EdTech)
A public registry of machine-readable learning-standard frameworks from all 50 US states and other issuing agencies, run by 1EdTech in the CASE JSON format. The digitally referenceable spine of what to teach, at which level and in what order, rather than the teaching content itself.
curriculumstandardsus - Legal & regulatoryOpen
Caselaw Access Project
6.9M US court decisions spanning the 1600s to 2020, digitised by Harvard Law School Library and released in a consistent machine-readable format. Covers every official, book-published state and federal case, with a free bulk API.
case-lawuspublic-domain - Multilingual & regional corporaOpen
CC-100
A high-quality multilingual corpus derived from Common Crawl. It splits every document into separate paragraphs, so it is effectively a paragraph-level rather than document-level corpus, which matters for your chunking strategy.
multilingualcommon-crawlparagraph-level - News, events & mediaOpen
CC-News
A subset of Common Crawl focused on news, containing millions of articles pulled from news websites around the world. It has fed several large language model training pipelines and gives you a ready news corpus without crawling sites yourself.
newsweb-crawlpretraining - Chemistry, materials & life sciencesOpen
ChEBI
A dictionary of small chemical compound molecular entities, with structure files and ontology files available for download. Particularly useful as a controlled vocabulary for chemical entity linking, where you match a mention in text to a standard identifier.
chemistryontologycontrolled-vocabulary - Chemistry, materials & life sciencesOpen
ChEMBL
A manually curated database of bioactive molecules with drug-like properties, bringing together chemical, bioactivity, and genomic data to support drug discovery. Holds close to 2.5M compound records on nearly 2M unique chemical structures, extracted mainly from the medicinal chemistry literature.
chemistrydrug-discoverybioactivity - Education & open learningOpen
CK-12 Foundation
Free K-12 library of customisable FlexBook textbooks, strong in maths and science, with adaptive practice, interactive simulations and study guides, organised by grade level and aligned to state standards. Non-commercial licence.
k-12textbooksmathematics - Biomedical & healthOpen
ClinicalTrials.gov
A US National Institutes of Health database of privately and publicly funded clinical studies. It holds structured records on more than 450,000 studies, including protocols, conditions, interventions, outcomes, and results, all in the public domain.
clinical-trialsuspublic-domain - Pre-embedded & RAG-readyOpen
Cohere Wikipedia 22-12 Embeddings (per-language)
Wikipedia encoded with the Cohere multilingual-22-12 embedding model, with embeddings (numeric vectors that capture meaning) computed on each article's title plus text. Published per language, covering many including Arabic, German, English, Spanish, French, Hindi, Italian, Japanese, Korean, Simple English, and Chinese. You can stream it rather than download it in full, which matters given the size.
wikipediaembeddingsmultilingual - Pre-embedded & RAG-readyOpen
Cohere Wikipedia Multilingual Embeddings (2023-11)
The November 2023 Wikipedia dump across 300+ languages, split into passages and turned into embeddings (numeric vectors that capture meaning) with the Cohere Embed V3 multilingual model. Close to 250 million paragraph embeddings in total. Drop it into a vector database for semantic search over all of Wikipedia, or use it as a knowledge source for multilingual RAG, including cross-lingual search where a query in one language returns relevant results from another.
wikipediaembeddingsmultilingual - Retrieval benchmarks & evaluationOpen
CoIR (Code Retrieval Benchmark)
A code retrieval benchmark spanning diverse programming tasks and languages. Built for evaluating how well models retrieve relevant source code, which matters if you are building RAG over codebases.
retrievalbenchmarkcode - Education & open learningOpen
Common Core State Standards
The US Common Core State Standards for English language arts and mathematics, defined grade by grade from kindergarten to grade twelve. A widely adopted curriculum framework, available in machine-readable form through ASN and CASE.
curriculumstandardsk-12 - Web corporaOpen
Common Corpus
An open, multilingual pretraining corpus of roughly 2.27 trillion tokens, built only from public-domain and permissively licensed text with documented provenance for every document. Maintained by PleIAs.
multilingualpretrainingpublic-domain - Web corporaOpen
Common Crawl
A nonprofit that crawls the web and freely provides its archives and datasets. Petabytes of raw web data from billions of pages, with a new snapshot each month. Used in the training of GPT-3, LLaMA, T5, and many other large language models.
web-crawlenglishmultilingual - Cultural heritage & archivesOpen
Common European Data Space for Cultural Heritage
The umbrella project coordinating cultural heritage data across Europe. Led by the Europeana Initiative, it helps cultural institutions and EU Member States share their collections through a common, shared infrastructure.
cultural-heritageeuropeinfrastructure - Web corporaOpen
Common Pile v0.1
An 8 TB corpus of public domain and openly licensed text drawn from 30 sources, assembled by EleutherAI as a licence-transparent alternative to web-scraped pretraining data. Used to train the Comma family of language models.
pretrainingpublic-domainopenly-licensed - Government & public sectorOpen
Companies House
The UK's official register of companies. Search every registered company, its directors, filing history, and accounts, or pull the data in bulk. A reliable source for UK corporate facts in a RAG system.
ukgovernmentcompanies - Knowledge graphs & structured dataOpen
ConceptNet
A multilingual common-sense knowledge graph that links words and phrases with labelled connections, such as "a cat is a pet" or "rain causes wet". It captures the everyday relationships between ideas that plain text rarely spells out.
knowledge-graphcommon-sensemultilingual - Multimodal & image-textLimited
Conceptual Captions (CC3M / CC12M)
A cleaned image alt-text dataset for automatic image captioning, produced by Google Research. Alt-text pulled from the web is filtered and generalised (hypernymed), replacing specific names with broader categories. Available in 3.3M (CC3M) and 12M (CC12M) variants.
image-textcaptioningalt-text - Climate & earth observationOpen
Copernicus (EU Earth Observation)
The EU's flagship Earth observation programme. The Copernicus Climate Change Service portal provides authoritative information about past, present, and future global climate, derived from Sentinel reference products alongside other satellite observations and in-situ measurements.
earth-observationsatelliteclimate - Academic & scientific literatureOpen
CORE
An aggregator of open-access research papers that harvests from thousands of repositories and journals worldwide. It holds over 300 million metadata records and more than 40 million full-text articles, all reachable through one search API.
open-accessaggregatorfull-text - RAG-specific & evaluationOpen
CRAG (Comprehensive RAG Benchmark)
Meta's factual question answering benchmark of 4,409 QA pairs across five domains, paired with mock web-search and knowledge-graph retrieval APIs. It is distinct from "Corrective RAG (CRAG)", which is a retrieval method, not a dataset.
rag-evaluationquestion-answeringbenchmark - Academic & scientific literatureOpen
Crossref
A nonprofit DOI registration agency that publishes metadata for over 150 million scholarly works, including journal articles, books, and conference proceedings. A DOI is the permanent identifier assigned to each work, and Crossref shares its metadata through a free API.
metadatadoicitations - Chemistry, materials & life sciencesOpen
Crystallography Open Database
A collection of over 350,000 crystal structure files covering organic, inorganic, and metal-organic compounds, released into the public domain under CC0.
crystallographymaterials-sciencepublic-domain - Multilingual & regional corporaOpen
CulturaX
A multilingual text dataset of 6.3 T tokens across 167 languages, built for training large language models. It combines cleaned versions of mC4 and OSCAR, two large web corpora, after a multi-stage filtering pipeline.
multilingualpretrainingcleaned - Government & public sectorOpen
data.europa.eu
The European Data Portal, a single point of access to open data from 36 European countries and the EU institutions. Over 1.6 million datasets span every policy area, giving RAG systems broad, multilingual coverage of official European data.
eueuropegovernment - Government & public sectorOpen
data.gov
The US federal government's open data portal, gathering more than 300,000 datasets from federal agencies. Coverage spans agriculture, climate, education, energy, finance, health, and public safety, making it a broad source of official US data for RAG systems.
usgovernmentopen-data - Government & public sectorOpen
data.gov.uk
The UK Government's central catalogue of open public sector data, run by the Government Digital Service. More than 47,000 datasets cover health, housing, transport, demographics, and the environment, giving RAG systems a trusted base of UK facts, figures, and official records.
ukgovernmentopen-data - Government & public sectorOpen
data.police.uk
Crime and policing data published on behalf of police forces in England and Wales, Northern Ireland, and the British Transport Police, gathered in one central place. Covers street-level crime, outcomes, and stop-and-search records.
ukcrimepolicing - Multimodal & image-textOpen
DataComp
A benchmark and dataset collection for training multimodal models. It provides a shared pool of image-text candidates and an evaluation framework so teams can test different data-curation strategies against a common yardstick.
image-textbenchmarkdata-curation - Data platforms & marketplacesOpen
Datahub.io
A platform for publishing and finding open data packages: datasets bundled with consistent metadata in standardised formats. It hosts curated collections of widely used reference data, from country codes to exchange rates, ready to drop into a pipeline.
data-platformopen-datadata-packages - Curated lists & meta-resourcesOpen
DataKind UK Open Data Sets
A curated list of UK-focused open datasets from DataKind UK, covering government, health, crime, housing, and social data. A quick way into British public data when you are building a RAG system with a UK focus.
directoryukopen-data - Encyclopaedic & general knowledgeOpen
DBpedia
A knowledge graph built by pulling the structured parts of Wikipedia, mainly the infoboxes, into machine-readable data. It holds billions of facts about people, places, organisations, and more as RDF triples, small subject, predicate, object statements, which you can search with the SPARQL query language.
knowledge-graphstructuredwikipedia - Web corporaOpen
DCLM-Baseline (DataComp-LM)
A filtered English web dataset from the DataComp-LM benchmark project, produced by running model-based quality filtering over Common Crawl. Built to show which data-curation choices most improve language model training.
web-crawlenglishpretraining - Code & technical documentationOpen
DevDocs
An open-source aggregator that pulls API documentation for major programming languages, frameworks, and tools into one searchable place. A tidy RAG source when you want up-to-date developer reference material without scraping dozens of separate documentation sites yourself.
api-documentationdeveloperaggregator - Cultural heritage & archivesOpen
Digital Public Library of America (DPLA)
An aggregator that brings together content from America's museums, libraries, and archives. Its API offers metadata on individual items and on collections, and the whole repository is available to download as zipped JSON files.
cultural-heritageusmuseums - Web corporaOpen
Dolma
A 3T token open corpus from Allen AI (its name stands for "Data for Open Language Models' Appetite") combining web text, scientific papers, code, public-domain books, Reddit posts, and Wikipedia. Built to train the OLMo models.
web-crawlpretrainingmulti-source - Biomedical & healthLimited
DrugBank
A freely accessible database that combines detailed drug data with drug target information, covering more than 15,000 drug entries. It links chemistry, pharmacology, and biology in one place, which makes it a strong knowledge source for medical and pharmaceutical RAG.
drugspharmacologydrug-targets - Legal & regulatoryOpen
eCFR
An up-to-date electronic Code of Federal Regulations with a full bulk API. Continuously updated, unlike the annual print edition.
usregulationspublic-domain - Legal & regulatoryOpen
EDGAR (SEC Filings)
The US Securities and Exchange Commission's repository of corporate filings, including annual reports (10-K), quarterly reports (10-Q), current reports (8-K), and proxy statements. You can search the full text or pull filings in bulk through an API, which makes it a rich source for financial and regulatory RAG.
usfinancial-filingssec - Speech & audioLimited
Emilia
A large-scale, multilingual in-the-wild speech dataset for speech generation, starting at over 101,000 hours across six languages and extended to more than 216,000 hours in Emilia-Large. Built for training natural, spontaneous TTS and speech-understanding models.
speechtext-to-speechmultilingual - Patents & intellectual propertyLimited
EPO Espacenet & Open Patent Services
Free access to over 140M patent documents from the European Patent Office (EPO). Its Open Patent Services text analysis tools have made patent full texts much easier to reach, though the terms are open access rather than an open reuse licence.
patentseuropeepo - Web corporaOpen
Essential-Web v1.0
A 24-trillion-token web corpus (about 23.6 billion documents) from Essential AI where every document carries a twelve-category taxonomy covering topic, format, complexity and quality. The labels let you carve out domain or quality subsets with simple filters, without training your own classifiers.
web-crawlpretrainingmetadata - Legal & regulatoryOpen
EUR-Lex
The official portal for European Union law. It provides EU treaties, legislation, case law, and legislative proposals in 24 languages, with bulk download available. A dependable source for multilingual legal and regulatory RAG across the EU.
eulegislationcase-law - Cultural heritage & archivesOpen
Europeana
A platform for discovering cultural heritage collections across Europe, with multilingual access to over 60 million digitised items from around 4,000 institutions, including books, paintings, maps, manuscripts, and audiovisual and 3D media.
cultural-heritageeuropemultilingual - Statistics & economicsOpen
Eurostat
The statistical office of the European Union, providing official statistics for Europe. Covers the economy, population, environment, agriculture, trade, and more across EU member states, with harmonised figures that let you compare countries on a like-for-like basis.
statisticseueurope - RAG-specific & evaluationLimited
FinanceBench
A benchmark for evaluating LLMs and retrieval systems on questions about public financial filings. The full set holds 10,231 question, answer and evidence triplets; a 150 example sample is released openly.
financequestion-answeringevaluation - Legal & regulatoryOpen
Find Case Law (UK National Archives)
Official UK court decisions published by the National Archives, with structured XML and a documented API. The authoritative modern source for downloading UK judgments.
case-lawukcourt-opinions - Web corporaOpen
FinePDFs
HuggingFace's corpus built entirely from PDFs: roughly 3 trillion tokens across 475 million documents in 1,733 languages, drawn from Common Crawl and the open web. It captures dense long-form content (reports, manuals, papers) that HTML web corpora miss, and is the PDF counterpart to FineWeb.
pdfmultilingualpretraining - Web corporaOpen
FineWeb / FineWeb-Edu
A 15T token English web corpus distilled from Common Crawl by HuggingFace, filtered and deduplicated for language model training. FineWeb-Edu is a subset filtered for educational content. One of the higher-quality open web datasets for pretraining and broad-coverage RAG.
web-crawlenglishpretraining - Web corporaOpen
FineWeb-Edu
The educational-quality subset of FineWeb, about 1.3 trillion English tokens kept by an educational-quality classifier, with a larger 5.4T token variant at a looser threshold. Built by HuggingFace as a denser, more teachable base for pretraining and broad-coverage RAG.
web-crawlenglishpretraining - Web corporaOpen
FineWeb2
A multilingual extension of FineWeb from HuggingFace, built with an open curation pipeline that adapts filtering and deduplication across languages. Useful when you need clean web text beyond English for pretraining or multilingual RAG.
web-crawlmultilingualpretraining - Multilingual & regional corporaOpen
FineWeb2-HQ
A high-quality multilingual pretraining corpus of roughly the top 10% of FineWeb2 documents in each of 20 languages, selected by a model-based quality classifier. Built by EPFL and released on Hugging Face.
multilingualpretrainingweb-crawl - Government & public sectorOpen
Fingertips (Public Health Data)
A large public health data collection for England, including mental health indicators and the Public Health Outcomes Framework. Useful for viewing trends over time and comparing differences between areas.
ukpublic-healthstatistics - RAG-specific & evaluationOpen
FlashRAG Benchmark Datasets
A research toolkit that bundles standardised, pre-processed benchmark datasets for RAG evaluation across many scenarios. Every dataset comes in a single unified JSONL format, so you can swap benchmarks in and out without rewriting your data loading each time.
evaluationbenchmarktoolkit - RAG-specific & evaluationOpen
FRAMES (Fact, Fetch, and Reason)
A Google evaluation set of 824 challenging questions that tests factuality, multi-hop retrieval and reasoning together, each paired with a gold answer and the Wikipedia articles needed to answer it. Built to measure end-to-end RAG.
rag-evaluationfactualitymulti-hop - Statistics & economicsOpen
FRED (Federal Reserve Economic Data)
A large online database of macroeconomic and financial time series maintained by the Federal Reserve Bank of St. Louis, with around 845,000 series drawn from over 120 public and private sources. Free to access through a REST API returning JSON or XML.
economicstime-seriesmacroeconomics - Legal & regulatoryOpen
Free Law Project / CourtListener
The largest open collection of US court opinions, oral argument recordings, and judicial financial disclosures, run by the nonprofit Free Law Project. Its RECAP Archive holds hundreds of millions of federal court docket entries, and journalists at the Wall Street Journal and ProPublica have used it for investigations. A solid base for US legal RAG.
legaluscase-law - Encyclopaedic & general knowledgeOpen
Freebase
A collaborative knowledge base once run by Google, now retired but still available as downloadable data dumps. It holds structured facts about millions of entities, and much of its content has since moved into Wikidata.
knowledge-basestructuredarchived - Retrieval benchmarks & evaluationOpen
FreshStack
A 2025 framework and benchmark for retrieval over fast-moving technical documentation. It pairs real Stack Overflow questions and answers with chunked GitHub code and docs across five niche domains, giving code and docs RAG a hard, contamination-resistant test.
retrievalbenchmarkcode - News, events & mediaOpen
GDELT
GDELT, the Global Database of Events, Language, and Tone, monitors print, broadcast, and web news in more than 100 languages from every country. It records events back to 1979, builds a Global Knowledge Graph, and adds sentiment and emotion analysis, all queryable on Google BigQuery.
newseventsmultilingual - Geospatial & mappingOpen
GeoNames
A geographical database of more than 11 million place names, each with coordinates, population, elevation, and administrative divisions, covering every country. It works well as a gazetteer for resolving place names to locations.
gazetteerplace-namescoordinates - Code & technical documentationOpen
GitHub Public Repositories (GH Archive / GHTorrent)
Two projects that capture public GitHub activity for analysis. GH Archive records the public GitHub event timeline as downloadable hourly archives, and GHTorrent offers a queryable mirror of GitHub metadata. Together they help you mine code, issues, pull requests, and documentation at scale.
githubevent-dataissues - Agentic & Tool-UseOpen
Glaive Function Calling v2
A widely used open dataset of about 113,000 synthetic multi-turn chat conversations that include function calls and their results, made by Glaive AI. One of the most downloaded open datasets for fine-tuning models to call tools, released under Apache 2.0.
function-callingtool-usesynthetic - Multilingual & regional corporaOpen
GlotCC
An open, broad-coverage corpus and processing pipeline built from Common Crawl, targeting minority and low-resource languages that larger corpora tend to miss. It fills gaps left by datasets focused on higher-resource languages.
multilinguallow-resourceminority-languages - Books & literatureOpen
Google Books Ngrams
Word and phrase frequency data drawn from Google's digitised book collection, spanning centuries of published text. It counts how often words and short phrases appear by year, which is useful for historical language analysis and time-aware RAG features.
ngramsword-frequencyhistorical - Data platforms & marketplacesOpen
Google Dataset Search
A search engine for datasets that indexes millions of them from thousands of repositories worldwide. It does not host data itself: it points you to wherever each dataset lives, which makes it a fast first stop when you are hunting for a source on a specific topic.
data-platformsearch-enginediscovery - Knowledge graphs & structured dataLimited
Google Knowledge Graph API
An API into Google's Knowledge Graph, the store of billions of facts about entities that powers the info panels you see beside search results. You send a name or query and get back matching entities with descriptions, types, and links. Free to use within rate limits.
apiknowledge-graphentities - Patents & intellectual propertyOpen
Google Patents Public Datasets
A query-based collection of over 120M patent documents from more than 100 patent offices worldwide, including applications, pre-grant publications, and granted patents. Accessible through Google BigQuery, with hundreds of millions of USPTO events also queryable.
patentsmultilingualbigquery - Pre-embedded & RAG-readyOpen
Google Satellite Embedding (AlphaEarth Foundations)
Global annual satellite embeddings produced by Google and Google DeepMind's AlphaEarth Foundations model. Every 10 metre pixel holds a 64 dimensional vector summarising a year of multi-sensor Earth observation, available yearly from 2017.
geospatialsatelliteembeddings - Geospatial & mappingOpen
Google-Microsoft-OSM Open Buildings (VIDA)
A conflated global building-footprint layer that merges Google Open Buildings V3, Microsoft Global ML Footprints, and OpenStreetMap into roughly 2.7 billion footprints, each labelled by its source. Built by VIDA and hosted on Source Cooperative in cloud-native formats, it offers more complete coverage than any single provider on its own.
building-footprintsgeospatialglobal - Legal & regulatoryOpen
GovInfo
US federal legislation, regulations, and congressional records, with bulk data and API access. The authoritative source for US federal government publications.
uslegislationregulations - Speech & audioOpen
Granary
An open multilingual speech corpus from NVIDIA covering roughly one million hours of audio across 25 European languages, built for automatic speech recognition and speech translation.
speechasrspeech-translation - Government & public sectorOpen
Hansard
The official word-for-word record of proceedings in the UK Parliament, both the Commons and the Lords. Full text searchable and available as XML and linked data going back centuries, it is a rich source of political and legislative debate for RAG systems.
ukparliamentgovernment - Patents & intellectual propertyOpen
Harvard USPTO Patent Dataset (HUPD)
A large-scale, structured corpus of US patent applications built specifically for machine learning and natural language processing research. It fills the gap left by mainstream patent search tools, which are not designed with the ML and NLP community in mind.
patentsususpto - Cultural heritage & archivesLimited
HathiTrust Digital Library
A partnership of academic and research institutions offering millions of digitised titles. Its Research Center provides computational access to the full corpus, including in-copyright works, under non-consumptive research terms.
digitised-booksacademiclibraries - Multilingual & regional corporaOpen
HPLT (High Performance Language Technologies)
An EU-funded project building open language resources from web crawls and the Internet Archive. Its MonoHPLT dataset contains monolingual text in 75 languages, released as openly as licensing allows.
multilingualweb-crawlinternet-archive - Data platforms & marketplacesOpen
Hugging Face Datasets
The largest open hub for machine learning datasets, with well over 100,000 datasets you can search, stream, and version. Many RAG-ready corpora live here, and its Python library lets you pull data straight into a pipeline in a few lines.
platformmachine-learningstreaming - Statistics & economicsOpen
IMF Data
A gateway to the IMF's global economic data, giving streamlined access to macroeconomic and financial statistics. Includes International Financial Statistics covering balance of payments, interest rates, national accounts, prices, production, trade, and population for more than 200 countries.
statisticseconomicsfinance - Books & literatureLimited
Institutional Books 1.0
A 242 billion token dataset of roughly 983,000 public-domain volumes digitised from Harvard Library's collections, spanning more than 250 languages, with both raw and post-processed OCR text plus rich bibliographic metadata.
public-domainbooksocr - News, events & mediaOpen
Internet Archive
A nonprofit digital library giving free access to millions of books, films, audio recordings, software, archived websites, and television news. Its Wayback Machine has saved more than 800 billion web pages, making it a deep well of historical and current text.
digital-librarynonprofitweb-archive - Data platforms & marketplacesOpen
Kaggle Datasets
A hub of user-contributed datasets spanning almost every domain, owned by Google. Alongside the data you get notebooks, competitions, and active discussion, which makes it a quick way to find something workable and see how others have already cleaned and used it.
data-platformcommunitynotebooks - Education & open learningOpen
Khan Academy
A free nonprofit learning platform organised as a subject-to-skill knowledge tree across maths, science, economics and the humanities, from primary through early college. Videos, articles and mastery-based exercises, but non-commercial licensing and no public data API.
educationk-12mastery-learning - Pre-embedded & RAG-readyOpen
LAION-400M with CLIP Embeddings
400 million image-text pairs filtered from Common Crawl using CLIP, a model that scores how well an image matches a caption, and released alongside their pre-computed CLIP embeddings and kNN indices for fast similarity search. The ready-made embeddings make it unusually easy to work with. Note that it ships image URLs, not the images themselves, so some links have decayed.
image-textclipembeddings - Multimodal & image-textOpen
LAION-5B
An image-text dataset of over 5.8B examples, built by filtering Common Crawl with a CLIP model that scores how well an image matches a caption. Includes 2.32B English pairs, 2.26B multilingual pairs, and 1.27B not tied to any particular language. Provides URLs, not the images themselves.
image-textclip-filteredweb-scale - Retrieval benchmarks & evaluationOpen
LegalBench-RAG
The first open benchmark for the retrieval step of legal RAG. It offers 6,858 human-annotated query-answer pairs over a corpus of roughly 79 million characters, with character-level ground-truth spans across contracts and privacy policies, assembled from CUAD, MAUD, ContractNLI and PrivacyQA.
retrievalbenchmarklegal - Legal & regulatoryOpen
Legislation.gov.uk
The official UK government legislation database, holding all UK Acts of Parliament and statutory instruments. It is available as XML and linked data, so you can load the full statute book into a RAG system rather than scraping individual pages.
uklegislationstatutes - Chemistry, materials & life sciencesOpen
LeMaterial (LeMat-Bulk)
A harmonised, standardised merge of three major computational materials databases (Materials Project, Alexandria and OQMD) into a single format of roughly 6.7 million entries with consistent property definitions.
materials-sciencechemistrycomputational - Education & open learningLimited
LibreTexts
One of the largest open textbook platforms, covering higher education with some K-12 material. Hosts interlinked textbooks and reference material across science, mathematics, humanities, and social sciences. Mostly NonCommercial, so commercial use is ruled out.
educationopen-textbookscc-by-nc-sa - Speech & audioOpen
LibriSpeech
A collection of 1,000 hours of read English speech, released under CC BY 4.0 and stored using the open-source FLAC audio encoder. Labels are aligned at the sentence level. Derived from LibriVox public domain audiobooks, it is the standard benchmark for English automatic speech recognition.
speechenglishasr - Books & literatureOpen
LibriVox
A volunteer project providing free public domain audiobooks. Volunteers record themselves reading works whose copyright has expired, which makes it a handy source of speech data for training speech-to-text models and building multimodal RAG.
public-domainaudiobooksaudio - Speech & audioOpen
LJ Speech
An open dataset of 13,100 short audio clips of a single speaker reading passages from seven non-fiction books. Every clip is transcribed and clips run from 1 to 10 seconds. Released into the public domain, it is the standard single-speaker text-to-speech benchmark.
speechenglishtext-to-speech - Government & public sectorOpen
London Datastore
A free and open data-sharing site for data and analysis from the Greater London Authority, providing more than 700 datasets at regional and local levels across the capital.
uklondonregional - Retrieval benchmarks & evaluationOpen
LongEmbed
A retrieval benchmark focused on very long documents, with an average length above 5,500 words. It addresses a gap in standard benchmarks, which mostly use short documents of at most 512 tokens (the chunks of text a model reads at once).
retrievalbenchmarklong-context - Multilingual & regional corporaOpen
MADLAD-400
A manually audited, general-domain monolingual dataset of 3 T tokens spanning 419 languages, built from Common Crawl by Google DeepMind and Google Research. It comes in a noisy version and a more heavily filtered clean version.
multilingualpretrainingcommon-crawl - Pre-embedded & RAG-readyOpen
Major TOM
An open, globally dense collection of embeddings and ML-ready imagery derived from Copernicus Sentinel data, published by ESA Phi-lab. The embedding expansions run to well over 100 billion vectors covering most of the Earth.
earth-observationsatelliteembeddings - Multilingual & regional corporaOpen
Masakhane
A grassroots research community building natural language processing resources for African languages. It produces datasets, benchmarks, and models across dozens of African languages that are otherwise poorly served.
multilingualafrican-languageslow-resource - Education & open learningOpen
Mason OER Metafinder
A discovery tool that runs a real-time simultaneous search across 21 different sources of open educational materials, rather than searching a static database. Useful for finding content to ingest, not a dataset in itself.
educationmeta-resourcesearch-tool - Retrieval benchmarks & evaluationOpen
Massive Legal Embedding Benchmark (MLEB)
A benchmark for evaluating legal text embedding and retrieval models, built from ten expert-annotated datasets spanning six jurisdictions and three task types (search, zero-shot classification, and question answering).
legalembeddingsretrieval - Chemistry, materials & life sciencesOpen
Materials Project
A database of computed properties for known and predicted materials, produced through high-throughput density functional theory calculations (a physics method for modelling how electrons behave in a material). Free API access with registration.
materials-sciencecomputed-propertiesapi - Mathematics & reasoningOpen
MegaMath
An open mathematics pretraining dataset curated from diverse, maths-focused sources, with over 300 billion tokens. Among the largest open maths corpora available.
mathematicspretraininglarge-corpus - Education & open learningOpen
MERLOT
A curated collection of online learning and support materials, plus content-creation tools, led by an international community of educators, learners, and researchers. Points to resources hosted elsewhere across many disciplines.
educationopen-educational-resourcescatalogue - Climate & earth observationOpen
Met Office DataPoint / CEDA
The Centre for Environmental Data Analysis hosts UK atmospheric and Earth observation data, including Met Office archives, climate model outputs, and satellite data.
climateweatherearth-observation - Biomedical & healthLimited
MIMIC-III / MIMIC-IV
Freely accessible critical care databases with de-identified health data from more than 40,000 intensive care patients at Beth Israel Deaconess Medical Center. They include vital signs, lab results, medications, and clinical notes. Access needs credentialing and a data-use agreement.
clinicalcritical-careehr - Retrieval benchmarks & evaluationOpen
MIRACL
A multilingual retrieval benchmark built from human-annotated relevance judgements over Wikipedia articles, covering 18 languages. Widely used for evaluating retrieval within a single language across a broad set of languages.
retrievalbenchmarkmultilingual - Biomedical & healthOpen
MIRIAD
A million-scale medical instruction and retrieval dataset of roughly 5.8 million question-answer pairs, each grounded in a passage from peer-reviewed biomedical literature. Built for medical RAG, retrieval, and hallucination detection.
biomedicalmedical-qainstruction-tuning - Education & open learningLimited
MIT OpenCourseWare
Teaching materials from more than 2,500 MIT undergraduate and graduate courses, published freely online. Includes lecture notes, assignments, and exams. Released under a NonCommercial licence, so it is free to use for non-commercial purposes only.
educationcoursewarelecture-notes - Web corporaOpen
MixtureVitae
A permissive-first, open pretraining corpus of roughly 422 billion tokens, drawn from public-domain, permissively licensed, and civic or government text, with sources organised into provenance-based risk tiers.
pretrainingpermissivepublic-domain - Speech & audioOpen
Mozilla Common Voice
A crowdsourced speech platform from Mozilla that releases datasets under the CC0 licence. Volunteers read sentences aloud and other community members validate each recording. As of release 19.0 it holds 32,584 hours of speech across 131 languages, making it one of the largest openly licensed voice datasets available.
speechmultilingualcrowdsourced - Retrieval benchmarks & evaluationLimited
MS MARCO
One of the most widely used general-domain retrieval benchmarks, built from real Bing search queries with human-annotated passage relevance judgements. It carries a non-commercial licence: some model developers deliberately exclude MS MARCO from training because of its terms, so check carefully before any commercial use.
retrievalbenchmarkquestion-answering - Retrieval benchmarks & evaluationOpen
MTEB (Massive Text Embedding Benchmark)
The Massive Text Embedding Benchmark, a common framework for comparing text embedding models (the numeric vectors that represent meaning for search and retrieval). It brings together SemEval, BEIR, and many other datasets across tasks such as retrieval, classification, and clustering. The MMTEB extension widens the scope to over 250 languages and more than 500 tasks.
embeddingsbenchmarkevaluation - Climate & earth observationOpen
NASA Earthdata
The home for full and open access to NASA's Earth science data collections. The umbrella portal for the largest collection of freely available Earth science data in the world, with over 12,400 datasets spanning land, ocean, atmosphere, and cryosphere, collected by satellite missions, airborne campaigns, and ground stations.
earth-observationsatelliteclimate - Geospatial & mappingOpen
Natural Earth
A public domain map dataset built for cartography, offered at three scales: 1:10m, 1:50m, and 1:110m. It bundles cultural data (borders, cities, roads), physical data (coastlines, rivers, lakes), and raster imagery, all ready to drop into a map.
public-domainmapscartography - Retrieval benchmarks & evaluationOpen
Natural Questions
A general-domain retrieval benchmark of real, anonymised Google search queries paired with Wikipedia passages that contain the answers. Alongside MS MARCO, one of the two most widely used retrieval benchmarks for general-domain evaluation.
retrievalbenchmarkquestion-answering - Mathematics & reasoningOpen
NaturalProofs
A dataset of 32,000 theorem statements and proofs, 14,000 definitions, and 2,000 other pages including axioms and corollaries, drawn from ProofWiki, the Stacks Project, and mathematics textbooks.
mathematicsproofstheorems - Knowledge graphs & structured dataOpen
NELL (Never-Ending Language Learner)
A machine learning system from Carnegie Mellon that has been reading the web since 2010 and building a knowledge base as it goes. It extracts beliefs, entities, and the relationships between them from text, and keeps refining them over time.
knowledge-graphmachine-learningweb-extraction - Multilingual & regional corporaOpen
Nemotron Post-Training Dataset v2
NVIDIA's 2025 post-training set of prompts and synthetic responses for supervised fine-tuning and reinforcement learning, covering mathematics, code, STEM, reasoning, and instruction following, with the instruction data expanded into five additional languages.
post-trainingfine-tuningsynthetic-data - Web corporaLimited
Nemotron-CC-v2
NVIDIA's rebuilt Common Crawl corpus for LLM pretraining, adding eight fresh 2024 to 2025 snapshots, synthetic rephrasing, and multilingual synthetic question and answer data, with mathematics and code content preserved.
web-crawlpretrainingcommon-crawl - Education & open learningOpen
Next Generation Science Standards
US K-12 science standards developed by a consortium of states and released in 2013. A curriculum framework of performance expectations arranged by grade band and disciplinary core idea, free to use with attribution and browsable online or downloadable as PDFs.
science-standardsk-12curriculum-framework - Government & public sectorOpen
NHS England Digital
The national provider of information and data on health and social care in England, covering general practices, hospital-level mortality, mental health, population health, and NHS outcomes. Published under the Open Government Licence.
ukhealthsocial-care - Curated lists & meta-resourcesOpen
NLP Datasets
An alphabetical list of free and public domain text datasets for natural language processing, covering corpora, dialogue, sentiment, and summarisation. Handy when you want text-heavy data to build or evaluate a RAG system.
awesome-listnlptext-corpora - Climate & earth observationOpen
NOAA Open Data Dissemination (NODD)
Free, full, and open access to NOAA data through partner platforms including Amazon Web Services, Google, and Microsoft Azure. Covers climate data records and atmospheric, oceanic, and terrestrial datasets from one of the largest climate archives in the world.
climateweatheroceanic - Government & public sectorOpen
Nomis
A service from the Office for National Statistics holding detailed data on the national and local labour market, covering official census and labour market statistics. Offers an API for programmatic access.
ukcensuslabour-market - Multimodal & image-textOpen
OBELICS
A web-scale dataset of 141M multimodal English web documents that interleave text and images, containing 353M images and 115 billion tokens. Unlike image-caption pair datasets, it keeps whole documents with text and images in their original order. Used to train the IDEFICS models.
image-textinterleavedweb-scale - Statistics & economicsOpen
OECD Data
Reports, data, and publications on the economic, environmental, and social conditions in hundreds of countries. Includes International Development Statistics covering aid volume, origin, and types, plus the Creditor Reporting System with detailed information on individual aid activities.
statisticseconomicsdevelopment - Education & open learningOpen
OER Commons
A digital library of over 500,000 open educational resources gathered from institutions around the world. Aggregates full courses, modules, lesson plans, videos, simulations, and assessments across every discipline and education level.
educationopen-educational-resourcescc-licensed - Speech & audioOpen
Omnilingual ASR Corpus
A Meta FAIR corpus of transcribed spontaneous speech covering roughly 348 under-served, low-resource languages, released in 2025 under CC BY 4.0. Around 3,350 hours of audio paired with human transcriptions.
speechtranscriptionmultilingual - Government & public sectorOpen
ONS (Office for National Statistics)
The UK's official statistics body. It publishes Census data, economic indicators, population estimates, labour market figures, and more, making it the authoritative source for UK statistics in a RAG system.
ukgovernmentstatistics - Consumer & product dataOpen
Open Beauty Facts / Open Products Facts
Sister projects to Open Food Facts that apply the same collaborative model to cosmetics (Open Beauty Facts) and general consumer goods (Open Products Facts). Contributors scan barcodes and photograph packaging to build an open database of ingredients and product information.
cosmeticsbeautyproducts - Consumer & product dataOpen
Open Food Facts
A collaborative, free, and open database of ingredients, nutrition facts, and information on food products from around the world. Contributors scan barcodes and photograph ingredient lists and nutrition tables, following the model of Wikipedia and OpenStreetMap. Covers millions of products across 140+ countries.
foodnutritionbarcode - Books & literatureOpen
Open Library
An Internet Archive project building a web page for every book ever published. It holds more than 20 million catalogue records and lends many titles digitally, making it a rich source of book metadata for RAG systems.
booksmetadatacatalogue - Chemistry, materials & life sciencesOpen
Open Materials 2024 (OMat24)
Meta FAIR's open dataset of more than 110 million DFT calculations of inorganic materials, sampled for structural and compositional diversity. Models trained on it top the MatBench-Discovery leaderboard for stability and formation-energy prediction.
materials-sciencechemistrydft - Chemistry, materials & life sciencesLimited
Open Molecules 2025 (OMol25)
A large-scale dataset of more than 100 million high-accuracy DFT calculations spanning 83 elements, built by Meta FAIR to train machine-learning interatomic potentials for molecular systems.
chemistrymolecular-dynamicsdft - RAG-specific & evaluationOpen
Open RAGBench (Vectara)
A RAG evaluation benchmark from Vectara built on 1,000 arXiv papers, with multimodal content extraction and 3,045 question-and-answer pairs across scientific domains. Designed to test retrieval and answer quality on real research documents rather than short, simplified passages.
evaluationbenchmarkarxiv - Chemistry, materials & life sciencesOpen
Open Reaction Database
An open-access chemical reaction database built to support machine learning for reaction prediction, synthesis planning, and experiment design. A collaborative effort spanning pharmaceutical companies, academia, and technology firms.
chemistryreactionsmachine-learning - Biomedical & healthOpen
Open Targets Platform
An open knowledge base of scored target-disease associations for drug discovery, integrating genetics, genomics, literature, and drug evidence from many public sources. Accessible via REST and GraphQL APIs, BigQuery, and bulk Parquet downloads.
biomedicaldrug-discoverygenomics - Education & open learningOpen
Open Textbook Library
A searchable catalogue of free, openly licensed textbooks, developed by the University of Minnesota. Open textbooks here are funded, published, and licensed to be freely used, adapted, and distributed, with faculty reviews on each entry.
educationopen-textbookshigher-education - Academic & scientific literatureOpen
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.
academicmetadatacitations - Consumer & product dataLimited
OpenCorporates
A global database of companies, their directors, and regulatory filings. The largest open database of companies in the world, though bulk and API access is commercially licensed rather than freely reusable.
companiescorporateregistry - Encyclopaedic & general knowledgeOpen
OpenCyc
The open release of Cyc, one of the oldest attempts to hand-build common-sense knowledge for machines. It holds hundreds of thousands of concepts and millions of assertions about how the everyday world fits together. Now archived, but the data is still available.
knowledge-baseontologycommon-sense - Education & open learningLimited
OpenLearn (Open University)
The Open University's free learning platform, offering courses and educational resources across a wide range of subjects and academic levels. A strong UK-based option for education content. Released under a NonCommercial licence, so commercial use is not permitted.
educationukcc-by-nc-sa - Consumer & product dataLimited
OpenSanctions
A database of sanctioned entities, politically exposed persons, and related risk data, aggregated from official sources worldwide. Links entities to OpenCorporates where both hold the same entity, so you can combine sanctions data with company control information.
sanctionscompliancerisk - Speech & audioOpen
OpenSLR
Open Speech and Language Resources, a hosting site for speech and language datasets, software, and models. It is the home of LibriSpeech and dozens of other language-specific speech corpora, making it a central catalogue for finding openly available voice data.
speechmultilingualasr - Legal & regulatoryOpen
OpenStates
An open-source platform tracking US state legislation in real time. Bulk data and API access for bills, legislators, and votes across all 50 states.
uslegislationopen-source - Education & open learningOpen
OpenStax
Openly licensed college textbooks from a non-profit initiative of Rice University. Peer-reviewed titles cover common undergraduate subjects, free to read online and available in low-cost print, and are used by millions of students each month.
textbookseducationopen-textbooks - Geospatial & mappingOpen
OpenStreetMap
A crowdsourced map of the world, built and maintained by a community of volunteers. It holds detailed geographic data on roads, buildings, land use, points of interest, and much more, which you can download as a full planet file or as smaller regional extracts.
crowdsourcedmapsgeographic - Multilingual & regional corporaLimited
OpenSubtitles
A large corpus of translated movie and TV subtitles. The latest version covers 60 languages with 2.6 billion sentences in total. It is valuable for conversational and colloquial language, and for parallel translation data.
multilingualsubtitlesparallel-corpus - Mathematics & reasoningOpen
OpenThoughts3-1.2M
An open reasoning dataset of 1.2 million rows (roughly 850k maths, 250k code and 100k science questions) each paired with a chain-of-thought answer. Released by the Open Thoughts team in June 2025, it is the training set behind OpenThinker3 and is licensed Apache 2.0.
reasoningchain-of-thoughtmathematics - Mathematics & reasoningOpen
OpenWebMath
A corpus of high-quality mathematical text extracted from the web. Filtered from over 200 billion HTML pages on Common Crawl down to 6.3 million documents totalling 14.7 billion tokens. Mostly mathematics, with substantial coverage of physics, computer science, and statistics.
mathematicsweb-crawlcommon-crawl - Government & public sectorOpen
Ordnance Survey OpenData
A portfolio of free geospatial datasets covering Great Britain from the national mapping agency, available for commercial and personal reuse under the Open Government Licence. Provides mapping, boundary, and location data through the OS Data Hub in formats including GeoPackage, GML, Shapefile, and CSV.
ukgeospatialmapping - Web corporaOpen
OSCAR
A multilingual web corpus (its name stands for Open Super-large Crawled Aggregated coRpus) extracted from Common Crawl with per-document language classification. Covers more than 150 languages, useful for non-English and cross-lingual RAG.
web-crawlmultilingualpretraining - Statistics & economicsOpen
Our World in Data
Research and interactive visualisations on global problems, built by aggregating and harmonising data from official sources. Every chart can be downloaded as a ZIP containing a CSV file, JSON metadata, and a README, and the underlying data is available programmatically via an API. Full provenance is documented for every dataset.
statisticsdata-visualisationglobal - Geospatial & mappingOpen
Overture Maps Foundation
An open map data project run under the Linux Foundation by a group of industry members. It combines and harmonises data from OpenStreetMap, Microsoft, Meta, and others into consistent, interoperable map layers you can build on freely.
mapsopen-datainteroperable - Patents & intellectual propertyOpen
PatentsView
A platform from the US Patent and Trademark Office (USPTO) for exploring US patent data, with visualisation and analysis tools. It connects patents, inventors, organisations, and locations in a regularly updated database, and offers both an API and complete bulk data tables.
patentsususpto - Multimodal & image-textOpen
PD12M (Public Domain 12M)
An image-text dataset built only from materials marked with a Public Domain Mark or released under Creative Commons Zero (CC0). 12.4M image-caption pairs, with a 3.3M subset called PD3M, sized to match the Conceptual Captions datasets while staying copyright-clean for commercial use.
image-textpublic-domaincc0 - Biomedical & healthOpen
PhysioNet
A repository of freely available medical research data, including physiological signals such as ECG and EEG, clinical databases, and related software. It hosts MIMIC and many other biomedical datasets, so it is a central hub for clinical RAG source material.
clinicalphysiological-signalsecg - Legal & regulatoryOpen
Pile of Law
A 256 GB dataset of open English-language legal and administrative text, covering court opinions, contracts, administrative rules, and legislative records. A ready starting point for building a legal RAG system without assembling sources yourself.
legalenglishus - RAG-specific & evaluationOpen
PleIAs RAG-Resources
A curated collection of datasets and resources assembled by PleIAs for building and evaluating RAG applications. A useful starting point when you want ready-made material to test retrieval pipelines without hunting for datasets one by one.
evaluationcollectioncurated - Books & literatureOpen
Project Gutenberg
A volunteer effort to digitise and archive public domain books, with more than 70,000 free ebooks. Mostly English, but it covers many languages. One of the oldest digital library projects, so it is a clean, permissively licensed source of full-text literature for RAG.
public-domainbooksliterature - Mathematics & reasoningOpen
Proof-Pile-2
A 55 billion token dataset of mathematical and scientific documents drawn from arXiv, OpenWebMath, and AlgebraicStack. The AlgebraicStack subset alone contributes 11 billion tokens covering numerical computing, computer algebra, and formal mathematics.
mathematicsarxivformal-mathematics - Chemistry, materials & life sciencesOpen
PubChem
An open chemistry database run by the US National Institutes of Health. Holds over 300M substances, 100M compounds, and almost 300M recorded bioactivities, covering chemical structures, identifiers, physical properties, patents, and health, safety, and toxicity data. Available in SDF, CSV, and other formats.
chemistrydrug-discoveryus-government - Academic & scientific literatureOpen
PubMed / PubMed Central (PMC)
The US National Library of Medicine's database of biomedical and life sciences literature. PubMed indexes over 36 million citations and abstracts, and PubMed Central (PMC) adds free full-text access to a growing subset. You can download the data in bulk over FTP.
biomedicalhealthfull-text - RAG-specific & evaluationOpen
RAG-Mini-Wikipedia
A small evaluation set that pairs 918 question-and-answer pairs with a matching corpus of 3,200 Wikipedia passages. Built specifically for testing RAG systems, it is small enough to run quick evaluation loops while still covering realistic retrieval questions.
evaluationquestion-answeringwikipedia - RAG-specific & evaluationOpen
RAGTruth
A word-level hallucination corpus of roughly 18,000 LLM responses generated in a RAG setting, each manually annotated for hallucinated spans across question answering, summarisation, and data-to-text tasks. A reference corpus for training and evaluating faithfulness detectors.
hallucination-detectionfaithfulnessbenchmark - Biomedical & healthOpen
RCSB Protein Data Bank
The canonical open archive of experimentally determined 3D biomolecular structures, over 250,000 of them, served alongside more than a million computed structure models. It is the ground truth of structural biology, released into the public domain under CC0.
proteinsstructural-biologybioinformatics - Web corporaOpen
RedPajama
An open reproduction of the LLaMA training data from Together. V1 aggregates Wikipedia, books, arXiv, GitHub, Common Crawl, and Stack Exchange; V2 is a web-only corpus of more than 100T tokens with 46 quality signals per document.
web-crawlpretrainingmulti-source - Code & technical documentationOpen
RefineCode (OpenCoder)
A reproducible code pretraining corpus of roughly 960 billion tokens spanning about 607 programming languages, built with a published cleaning pipeline. It is the pretraining data behind the OpenCoder code models.
codepretrainingreproducible - Web corporaOpen
RefinedWeb
A filtered English web dataset from the Technology Innovation Institute, creators of the Falcon models. It showed that carefully cleaned web-only data can match curated multi-source corpora for language model training.
web-crawlenglishpretraining - Government & public sectorOpen
Regulations.gov
The official US federal rulemaking portal, exposing a public API to dockets, proposed and final rules, supporting materials and millions of public comments across federal agencies. Free to access with a registered API key.
usgovernmentlegal - Multilingual & regional corporaOpen
ROOTS
The Responsible Open-Science Open-Collaboration Text Sources corpus, a 1.6 TB dataset spanning 46 natural languages and 13 programming languages, released by BigScience. Built from community-selected sources plus OSCAR, it trained BLOOM.
multilingualbigsciencebloom - Retrieval benchmarks & evaluationOpen
RTEB (Retrieval Embedding Benchmark)
The MTEB team's retrieval-focused embedding benchmark, launched in beta in October 2025 as a new retrieval section of the MTEB leaderboard. It spans 20 languages and enterprise domains such as law, healthcare, finance and code, and deliberately mixes open datasets with held-out private ones to measure genuine generalisation rather than training-set memorisation.
retrievalembeddingsbenchmark - Education & open learningOpen
Saylor Academy
A nonprofit offering more than 300 free, self-paced college-level online courses across many subjects, with curated readings, framing text, and assessments. Course outlines are CC BY licensed; embedded third-party materials keep their own varying licences.
educationopen-coursewarecurriculum-aligned - Statistics & economicsOpen
SDMX (Statistical Data and Metadata Exchange)
An open standard for exchanging statistical data and metadata electronically, developed jointly by the BIS, ECB, Eurostat, IMF, OECD, UN, and World Bank. Not a dataset itself but the interchange format many statistical agencies use to publish their figures.
standardstatisticsmetadata - Academic & scientific literatureOpen
Semantic Scholar / S2ORC
Allen AI's academic graph, covering hundreds of millions of papers linked by billions of citations. S2ORC, the Semantic Scholar Open Research Corpus, is the downloadable dataset with full text and parsed references, drawn from arXiv, PubMed, Crossref, publishers, and web crawlers.
academiccitationsfull-text - Education & open learningOpen
Siyavula
Free, CAPS-aligned Everything Maths and Everything Science textbooks from Siyavula, covering South African school grades. Read online or download as PDF and EPUB, openly licensed for copying and redistribution.
textbooksmathematicsphysical-sciences - Web corporaOpen
SlimPajama
A cleaned and deduplicated version of RedPajama-V1 from Cerebras. It drops short documents and removes near-duplicate content, cutting 1.2T tokens down to a denser 627B tokens without losing coverage.
web-crawlpretrainingcleaned - Cultural heritage & archivesOpen
Smithsonian Open Access
Millions of digital items from the Smithsonian's museums, research centres, and archives, released under CC0. Includes images, 3D models, and research datasets you can reuse freely, including commercially.
cultural-heritageusmuseums - Data platforms & marketplacesCommercial
Snowflake Marketplace
A data marketplace connecting more than 820 providers who offer over 3,400 live datasets, data services, and applications. Listings span financial, weather, demographic, and industry data, with a mix of free and paid options delivered straight into a Snowflake account.
data-platformmarketplacecommercial - Code & technical documentationOpen
Software Heritage
A non-profit universal archive built specifically for software source code, preserving not just files but their full development history. Over 27 billion unique source files collected from public repositories, each identified by a cryptographic hash. The upstream source for The Stack v2.
source-codearchivepreservation - Code & technical documentationOpen
Stack Exchange Data Dump
Periodic XML dumps of every site in the Stack Exchange network, including Stack Overflow. Each dump holds questions, answers, comments, votes, and user profiles. One of the richest question-and-answer datasets you can freely download, and a natural fit for building technical support and developer RAG systems.
question-answeringstack-overflowcommunity - Mathematics & reasoningOpen
StackMathQA
A curated collection of 2 million mathematical questions and answers sourced from various Stack Exchange sites. The question-and-answer structure maps naturally onto retrieval, which makes it well suited to Q&A-style RAG.
mathematicsquestion-answerstack-exchange - Books & literatureOpen
Standardized Project Gutenberg Corpus (SPGC)
A research-ready version of the Project Gutenberg catalogue with consistent formatting, tidy metadata, and token counts for every book. It saves you the cleanup work, so you get uniform full-text literature ready to chunk and embed for RAG.
public-domainbooksliterature - Encyclopaedic & general knowledgeOpen
Structured Wikipedia
Wikipedia rendered as pre-parsed, machine-readable JSON: abstracts, short descriptions, infoboxes, sections, parsed tables and references, with links to Wikidata entities. A beta from Wikimedia Enterprise covering nine languages, also mirrored on Hugging Face. The section-segmented shape a RAG pipeline actually wants.
encyclopaediastructuredjson - Encyclopaedic & general knowledgeOpen
SYNTH
A fully open synthetic corpus of amplified multilingual encyclopaedic text with built-in reasoning traces and exercises covering RAG, information extraction and QA. Released by PleIAs with the AI Alliance, it targets training and evaluating small, grounded, citeable reasoning models.
syntheticreasoning-tracesmultilingual - RAG-specific & evaluationOpen
TΒ²-RAGBench
A 2025 benchmark of 23,088 context-independent question, context and answer triples over more than 7,300 financial documents that mix text and tables. Each question maps to exactly one ground-truth document, making it purpose-built for numerical, table-aware financial RAG.
financequestion-answeringretrieval-benchmark - Speech & audioOpen
The People's Speech
A large, freely licensed English speech recognition dataset from MLCommons, assembled from openly licensed sources and totalling more than 30,000 hours. Built as a permissively licensed alternative for training commercial automatic speech recognition systems.
speechenglishasr - Web corporaOpen
The Pile
An 825 GB curated English text dataset from EleutherAI, made of 22 sub-datasets spanning books, academic papers, code, web content, and more. Built as a diverse training corpus and widely used to train early open language models.
pretrainingmulti-sourceenglish - Code & technical documentationOpen
The Stack v1 / v2
A large source code dataset from the BigCode project, built from permissively licensed code on GitHub with duplicate and near-duplicate files removed. Version 2 draws from Software Heritage and covers 600+ programming languages, making it a strong base for training and retrieval in code-aware RAG systems.
source-codepermissive-licencepretraining - Code & technical documentationOpen
The Stack v2
BigCode's large-scale source code corpus and the pretraining set behind StarCoder2, built with Software Heritage and spanning more than 600 programming languages. The full version runs to about 67.5TB, with records pointing to code held in Software Heritage's S3 rather than embedding it directly.
source-codepermissive-licencepretraining - Agentic & Tool-UseOpen
ToolBench
An open instruction-tuning dataset for teaching general tool use to language models, built for the ToolLLM project over 16,464 real-world REST APIs from RapidAPI across 49 categories, with single-tool and multi-tool, multi-step solution paths.
agentictool-usefunction-calling - Agentic & Tool-UseOpen
Toucan-1.5M
The largest open tool-agentic dataset: over 1.5 million trajectories synthesised from 495 real-world MCP servers spanning 2,000 plus tools, with multi-turn, sequential and parallel tool calls backed by real executions and error handling. Built by Agent-Ark and released under Apache 2.0, it is premier open data for training retrieval-and-tool (MCP) agents.
agentictool-usemcp - Web corporaOpen
TxT360
An open pre-training corpus of 15 trillion-plus tokens from LLM360, built by globally deduplicating 99 Common Crawl snapshots and blending in 14 curated domains such as FreeLaw, PG-19, Wikipedia, and scientific papers. Ships with a fully documented, reproducible processing recipe.
web-crawlpretrainingdeduplicated - Curated lists & meta-resourcesLimited
UK Data Service
The UK's largest collection of economic, population, and social research data for teaching, learning, and public benefit. Many datasets need registration or an institutional login, so plan for an access step before you build with them.
uksocial-scienceresearch - Education & open learningOpen
UK National Curriculum (England)
The statutory programmes of study and attainment targets for state schools in England, organised by key stage and subject and published by the Department for Education. The curriculum spine for England, freely reusable under the Open Government Licence.
curriculumukstandards - Statistics & economicsOpen
UN Data
A single access point to the UN system's statistical databases: 32 databases holding over 60 million records from 1970 onward. Covers greenhouse gas inventories, commodity trade statistics, energy statistics, gender, and key global development indicators.
statisticsunglobal - Biomedical & healthOpen
UniProt
A resource for protein sequence and functional information, with more than 250 million protein sequences and rich annotations. It is a core reference for bioinformatics, so it grounds life sciences RAG systems in reliable protein facts.
proteinssequencesbioinformatics - Academic & scientific literatureOpen
Unpaywall
A database that tracks where scholarly articles can be read legally for free, covering over 30 million open-access papers. Maintained by the nonprofit OurResearch, it links each article's DOI to full-text versions hosted across repositories and journals.
open-accessdoimetadata - Pre-embedded & RAG-readyOpen
Upstash Wikipedia 2024 BGE-M3 Embeddings
The June 2024 Wikipedia dump split into paragraphs and pre-embedded with the multilingual BGE-M3 model, roughly 144 million vectors across the 11 most popular languages. Each paragraph is prefixed with its article title before embedding and very short paragraphs are dropped, so you can skip the embedding step and load meaning-ready vectors straight into a vector database.
wikipediaembeddingsbge-m3 - Retrieval benchmarks & evaluationOpen
ViDoRe (Visual Document Retrieval Benchmark)
A benchmark for OCR-free, page-image document retrieval, introduced with the ColPali paper. It bundles page-level retrieval tasks across several domains and languages so vision-based retrievers can be evaluated directly on document images.
retrieval-benchmarkdocument-retrievalmultimodal - Speech & audioLimited
VoxCeleb
An audio dataset of more than 100,000 utterances from over 7,000 speakers, extracted from interview videos uploaded to YouTube. Widely used for speaker recognition and verification. The annotations are CC BY 4.0, but the terms of the underlying source videos still apply.
speechspeaker-recognitionyoutube - Climate & earth observationOpen
WeatherBench 2
Google Research's open benchmark and curated ERA5-derived data archive for evaluating data-driven global weather models over the one-to-fifteen-day range. It bundles an Apache-licensed evaluation framework, cloud-hosted ground-truth and baseline forecast datasets in Zarr, and a continuously updated leaderboard ranking models such as GraphCast and Pangu-Weather against ECMWF's IFS.
weather-forecastingbenchmarkera5 - Encyclopaedic & general knowledgeOpen
Wikidata
A free, collaborative knowledge base that holds structured facts for Wikipedia and the other Wikimedia projects. More than 100M items, each one machine-readable with statements, relationships, and identifiers, so you can pull clean facts instead of parsing article prose. Released into the public domain under CC0.
knowledge-basestructuredmultilingual - Pre-embedded & RAG-readyOpen
Wikidata Embedding Project
A hosted semantic-search service over Wikidata from Wikimedia Deutschland, built with Jina.AI and DataStax, serving vector embeddings of nearly 120 million items. You query it by meaning through a REST API and a native Model Context Protocol endpoint rather than downloading files, with reranking and structured context returned for each hit.
wikidataembeddingsknowledge-graph - Cultural heritage & archivesOpen
Wikimedia Commons
A media repository that hosts most of the images, video, audio, and other files used across Wikimedia projects. Over 100 million freely licensed files, with structured metadata through Structured Data on Commons, usable by anyone for almost any purpose.
mediaimagesfreely-licensed - Encyclopaedic & general knowledgeOpen
Wikipedia
Complete database dumps of Wikipedia and the other Wikimedia projects, in every language, refreshed roughly twice a month. The most widely used knowledge source for RAG systems, and an easy first corpus to build a retrieval pipeline on.
encyclopaediamultilingualgeneral-knowledge - Education & open learningOpen
Wikiversity
Free, collaboratively written learning materials from the Wikimedia Foundation: lessons, courses and tutorials organised by subject and educational level, in many languages. Available as bulk dumps and via API, like the other Wikimedia projects.
open-educational-resourceswikimediacc-by-sa - Patents & intellectual propertyOpen
WIPO PATENTSCOPE
A search service for international patent collections, maintained by the World Intellectual Property Organization (WIPO). It lets you search across patent applications filed under international treaties alongside national collections from many countries.
patentsglobalmultilingual - Multimodal & image-textOpen
WIT (Wikipedia-based Image Text)
A Wikipedia-based image-text dataset for multimodal, multilingual machine learning. It pairs images with their Wikipedia captions and the surrounding article text across many languages, giving richer context than a single caption.
image-textmultilingualwikipedia - Knowledge graphs & structured dataOpen
WordNet
A lexical database of English that groups nouns, verbs, adjectives, and adverbs into sets of synonyms called synsets, then links those sets by meaning. It maps how words relate, which sense means what, what is a kind of what, so software can work with meaning rather than just spelling.
lexical-databaseenglishsynonyms - Statistics & economicsOpen
World Bank Open Data
Free and open access to global development data from the World Bank, including the World Development Indicators: more than 900 indicators with time series for 210 countries from 1960 to the present. Also includes a Microdata Library of sample survey data and a Data Catalog for bulk downloads.
developmenteconomicsstatistics - Climate & earth observationOpen
World Resources Institute
Free environmental data for GIS, climate research, and sustainability projects. Includes Global Forest Watch and other environmental monitoring datasets.
environmentclimategis - Agentic & Tool-UseLimited
xLAM Function-Calling (APIGen)
Salesforce's APIGen-generated dataset of 60,000 verified function-calling examples spanning 3,673 executable APIs, each checked through format, execution and semantic stages. The training data behind the xLAM action models and core capability data for tool-using agents.
agentictool-usefunction-calling - Encyclopaedic & general knowledgeOpen
YAGO
A semantic knowledge base that combines facts from Wikipedia, WordNet, and GeoNames into a single, high-accuracy collection of statements about entities. It adds when and where each fact holds, so you get temporal and spatial detail alongside the plain relationships.
knowledge-basestructuredwikipedia - Multimodal & image-textOpen
YFCC100M
A multimedia research dataset of 100M Flickr photos and videos with their metadata, all published under Creative Commons licences. The exact licence varies from one item to the next, so check each one before you reuse it.
image-textvideoflickr - Speech & audioOpen
YODAS2
YODAS2 is the long-form edition of ESPnet's YODAS corpus, offering over 500,000 hours of Creative Commons YouTube speech across 149 languages, re-released as full-video audio at 24 kHz. It carries both labelled subsets (manual and automatic subtitles) and unlabelled audio for self-supervised work.
speechmultilingualaudio - Data platforms & marketplacesOpen
Zenodo
A general-purpose open repository built by CERN where researchers can deposit datasets, software, reports, and any other digital output. Every upload gets a DOI, a permanent identifier that makes the work easy to cite and find again, which makes Zenodo a reliable long-term home for research data.
data-platformopen-accessresearch - Web corporaOpen
Zyda-2
Zyphra's five trillion token English pretraining mixture, built by filtering and cross-deduplicating DCLM, FineWeb-Edu, Zyda-1 and Dolma's Common Crawl portion. A ready high-quality drop-in base that Zyphra reports beats its component datasets on downstream evaluation. Published on HuggingFace as Parquet under an open licence.
web-crawlenglishpretraining