Awesome LegalTech
Awesome LegalTech is a community-maintained directory, kept in a single GitHub repository by Vaquill AI, that catalogues the wider legal technology landscape rather than datasets alone. It lists open-source platforms, AI models, MCP servers, companies, tools, and datasets from across the global legal ecosystem, each as a short entry with a link out. The breadth is the point: it is a survey of who and what is active in legal tech, useful for orientation as much as for finding a specific resource.
You use it by browsing the README on GitHub and following the links that match what you are building. Because legal technology moves quickly, a living list like this is a practical way to keep track of new entrants, and the repository's activity (recent commits, merged pull requests) tells you whether it is being kept current. Starring or watching it is an easy way to catch additions over time.
For a RAG builder, the honest framing is that this is a discovery aid, not a data source. Its datasets section can point you towards corpora worth investigating, but much of the list is tools, products, and companies that will not feed a retrieval index directly. Treat it as reconnaissance: a way to understand the space, spot the players, and find leads, which you then evaluate on their own terms. If your goal is specifically legal datasets, the companion Awesome Legal Data list is the more focused starting point.
The caveats are those of any curated list. Coverage is uneven, inclusion is not an endorsement, and links go stale, so verify anything before you rely on it. Nothing here is licensed as a bundle: each linked project and dataset sets its own terms, and legal material in particular is often free to read but restricted to reuse, so check the licence at the source before you build on it commercially.
Related sources
AustLII
A free resource of medium-neutral case law and unreported judgments for all Australian jurisdictions, covering the whole country since 1995.
BAILII (British and Irish Legal Information Institute)
Free access to British and Irish primary legal materials, covering UK and Ireland case law and legislation. Alongside the National Archives, one of the main free sources for reading UK judgments.
Cambridge Law Corpus
A research dataset of more than 250,000 UK court cases, mostly from the 21st century but with some reaching back to the 16th century. Built by the University of Cambridge for legal natural language processing work, it is available under restricted access for research use.
CanLII
A national legal information institute covering Canada's federal and provincial jurisdictions, hosting over 300 databases of legislation and case law. Bilingual, with thousands of commentaries on Canadian court decisions.