Skip to content
RAG Repo

Awesome Public Datasets

Awesome Public Datasets is one of the longest-running and best-known "awesome" lists on GitHub, a community-curated index of high-quality open datasets organised under clear topic headings: agriculture, biology, climate, economics, education, finance, government, healthcare, and many more. It does not host data itself; each entry is a short description and a link to the primary source, which fits neatly with the principle of always going to the official home of a dataset rather than a mirror.

You use it as a discovery tool. Open the repository, jump to the topic that matches your project, and work through the links to find candidate datasets. Because the list is broad and cross-domain, it is especially handy when you know the subject area you need (say, climate or public health) but not yet the specific dataset, and it often surfaces sources you would not have thought to search for.

For RAG, this is a map rather than an ingredient. It shines at the survey stage, when you are scoping what open data exists for a domain, and it can quickly turn a vague need into a shortlist of real sources to evaluate. Once you have candidates, you follow the links and assess each one properly for coverage, format, and freshness.

The caveats are the familiar ones for community lists maintained through pull requests. Coverage is broad but uneven, quality varies between entries, and some links age faster than others, so treat inclusion as a lead to verify, not a guarantee. Critically, the list has no single licence: every dataset it points to sets its own terms, and those differ widely, so always check the licence and how current the data is at the source before you build on it. Alongside more specialised indices like Awesome Legal Data or the maths lists in this directory, Awesome Public Datasets is the broadest general-purpose starting point of the lot.

awesome-listdirectoryopen-datacommunitygithub

Related sources