Skip to content
RAG Repo

How we review sources

RAG Repo is a curated directory, not an automatically scraped list. Every source is chosen, described, and checked by hand against a consistent set of standards. This page explains that method, so you can judge how much to trust what you read here, and hold us to it.

What earns a place in the directory

We favour accuracy over completeness. We would rather describe a smaller number of sources well than list a huge number badly. A data source earns an entry when it meets three tests:

We always link to the primary source, the official site or repository, never a mirror or a third-party reseller. If a source is discontinued but still valuable as an archive, we mark it as archived rather than quietly dropping it.

How we describe each source

Every entry follows the same structure, so you can compare sources at a glance. For each one we record a plain-English description of what it is and what it is good for, its licence, its approximate size, the formats it ships in, how often it updates, who maintains it, and links to the documentation, downloads, API, and code where they exist. The descriptions are written to be neutral: we tell you what a source is and where it fits, and let you decide, rather than ranking or hyping.

A machine-readable catalogue

Everything in the directory is also published as a single JSON file at/sources.json, regenerated on every build. For each source it gives the category, formats, access patterns (how you actually fetch the data), a coarse size tier, the licence and our licence-clarity flags, and links back to the primary source and its page here. It carries a legend that documents every field. Reuse it with attribution to RAG Repo; each source keeps its own licence, recorded in the file.

How we judge access

Every source carries one of three access tiers, shown as a coloured badge, so you can see the practical barrier to entry before you click through:

How we assess RAG-readiness

Two sources can both be excellent and still take very different amounts of work to use. To capture that, we tag each one with how ready it is to drop into a retrieval pipeline:

This is an editorial judgement about the data as distributed, meant to set expectations, not a precise technical grade. You can read more about these stages in our Learn section.

How we handle licences

Licensing is where directories most often mislead people, so we treat it carefully. We state the licence on every entry. Where a licence is genuinely unclear, we say so plainly rather than guess. Beyond the licence name, we record four practical flags that answer the questions teams actually ask before shipping:

These flags are our reading of the published terms, provided to help you shortlist. They are not legal advice. Before you build on any source commercially, confirm the terms with the licence text yourself, and take proper advice if anything material rests on it. Our guide to choosing a dataset licence goes into this in more depth.

Why money never influences the directory

RAG Repo is a community resource that also needs to be sustainable, so it does carry advertising and, in clearly labelled blog posts, may carry affiliate links. We keep a firm wall between that and the directory. Source listings, their ordering, which sources are featured, and the editorial notes are never affected by advertising or any commercial relationship. There are no sponsored listings, and there never will be. Where money appears, it is a single, plain ad unit on a page, or a disclosed affiliate link inside an opinion piece, never a thumb on the scale of what we recommend.

Keeping it current, and fixing mistakes

Data sources change: licences are revised, datasets are archived, links move. We review entries over time, but we also rely on readers. If you spot something out of date or wrong, or know a source we should include, please tell us on thecontribute page. Corrections are welcome and we act on them, because a reference is only as good as its accuracy.