Skip to content
RAG Repo

Find data for your RAG AI system.

A curated directory of open datasets, knowledge bases, and data repositories for AI. Ideal for Retrieval-Augmented Generation (RAG), fine-tuning, and any project that runs on high-quality data you can trust.

177

data sources

27

categories

158

open access

Browse by category

View all sources
11 sources

Web corpora

Large-scale web-crawled text datasets

6 sources

Encyclopaedic & general knowledge

Wikipedia, Wikidata, and structured knowledge bases

7 sources

Academic & scientific literature

Research papers, preprints, and citation graphs

4 sources

Knowledge graphs & structured data

Semantic knowledge graphs and ontologies

5 sources

Code & technical documentation

Source code repositories and developer Q&A

15 sources

Legal & regulatory

Court opinions, legislation, and regulatory filings

12 sources

Government & public sector

Official government open data portals

4 sources

Geospatial & mapping

Geographic data, maps, and place databases

3 sources

News, events & media

News archives, event databases, and media datasets

5 sources

Books & literature

Public domain books and literary corpora

5 sources

Biomedical & health

Clinical data, drug databases, and medical literature

8 sources

Data platforms & marketplaces

Platforms for discovering and trading datasets

4 sources

RAG-specific & evaluation

Benchmark datasets designed for RAG evaluation

5 sources

Curated lists & meta-resources

Awesome-lists and dataset directories

9 sources

Multilingual & regional corpora

Large text datasets beyond English

6 sources

Speech & audio

Spoken-language and audio datasets

7 sources

Multimodal & image-text

Paired image and text datasets

8 sources

Mathematics & reasoning

Maths, proofs, and technical reasoning

8 sources

Retrieval benchmarks & evaluation

Standard benchmarks for retrieval quality

3 sources

Pre-embedded & RAG-ready

Datasets that arrive already vectorised

5 sources

Patents & intellectual property

Patent full text and metadata

6 sources

Cultural heritage & archives

Museum, library, and archive collections

7 sources

Chemistry, materials & life sciences

Structured scientific and chemical data

5 sources

Climate & earth observation

Satellite, weather, and environmental data

7 sources

Statistics & economics

Official statistics and economic indicators

8 sources

Education & open learning

Open textbooks and course materials

4 sources

Consumer & product data

Product, food, and company databases

Featured sources

A good place to start, whatever you are building.

Pre-embedded & RAG-readyOpen

Cohere Wikipedia Multilingual Embeddings (2023-11)

The November 2023 Wikipedia dump across 300+ languages, split into passages and turned into embeddings (numeric vectors that capture meaning) with the Cohere Embed V3 multilingual model. Close to 250 million paragraph embeddings in total. Drop it into a vector database for semantic search over all of Wikipedia, or use it as a knowledge source for multilingual RAG, including cross-lingual search where a query in one language returns relevant results from another.

wikipediaembeddingsmultilingual