Skip to content
RAG Repo

Find data for your RAG AI system.

A curated directory of open datasets, knowledge bases, and data repositories for AI. Ideal for Retrieval-Augmented Generation (RAG), fine-tuning, and any project that runs on high-quality data you can trust.

238

data sources

28

categories

213

open access

Browse by category

View all sources
20 sources

Web corpora

Large-scale web-crawled text datasets

8 sources

Encyclopaedic & general knowledge

Wikipedia, Wikidata, and structured knowledge bases

7 sources

Academic & scientific literature

Research papers, preprints, and citation graphs

4 sources

Knowledge graphs & structured data

Semantic knowledge graphs and ontologies

7 sources

Code & technical documentation

Source code repositories and developer Q&A

15 sources

Legal & regulatory

Court opinions, legislation, and regulatory filings

13 sources

Government & public sector

Official government open data portals

5 sources

Geospatial & mapping

Geographic data, maps, and place databases

3 sources

News, events & media

News archives, event databases, and media datasets

6 sources

Books & literature

Public domain books and literary corpora

9 sources

Biomedical & health

Clinical data, drug databases, and medical literature

8 sources

Data platforms & marketplaces

Platforms for discovering and trading datasets

9 sources

RAG-specific & evaluation

Benchmark datasets designed for RAG evaluation

5 sources

Curated lists & meta-resources

Awesome-lists and dataset directories

11 sources

Multilingual & regional corpora

Large text datasets beyond English

10 sources

Speech & audio

Spoken-language and audio datasets

7 sources

Multimodal & image-text

Paired image and text datasets

9 sources

Mathematics & reasoning

Maths, proofs, and technical reasoning

13 sources

Retrieval benchmarks & evaluation

Standard benchmarks for retrieval quality

7 sources

Pre-embedded & RAG-ready

Datasets that arrive already vectorised

6 sources

Patents & intellectual property

Patent full text and metadata

6 sources

Cultural heritage & archives

Museum, library, and archive collections

10 sources

Chemistry, materials & life sciences

Structured scientific and chemical data

6 sources

Climate & earth observation

Satellite, weather, and environmental data

8 sources

Statistics & economics

Official statistics and economic indicators

17 sources

Education & open learning

Open textbooks and course materials

4 sources

Consumer & product data

Product, food, and company databases

5 sources

Agentic & Tool-Use

Tool-calling, function-calling, and agent trajectory datasets

Featured sources

A good place to start, whatever you are building.

Pre-embedded & RAG-readyOpen

Cohere Wikipedia Multilingual Embeddings (2023-11)

The November 2023 Wikipedia dump across 300+ languages, split into passages and turned into embeddings (numeric vectors that capture meaning) with the Cohere Embed V3 multilingual model. Close to 250 million paragraph embeddings in total. Drop it into a vector database for semantic search over all of Wikipedia, or use it as a knowledge source for multilingual RAG, including cross-lingual search where a query in one language returns relevant results from another.

wikipediaembeddingsmultilingual

Read and learn

More than a directory: guides, comparisons, and research to help you choose and use data well.

From the blog

All posts →

Practical guides, comparisons, and roundups for building with data.

From the research

All research →

Deeper reading on where retrieval is heading, and what it means today.