EUR-Lex
European Union law all lives in one place here, and the coverage is genuinely complete: the founding treaties, every regulation and directive in force, consolidated texts that merge an act with its later amendments, rulings from the Court of Justice, bills still moving through the legislative pipeline, and the Official Journal stretching back to the 1950s. Every item carries a CELEX number, subject classifications, and typed links to the acts it amends, repeals, or cites.
Three access routes matter for building a corpus. You can pull individual documents as XML or HTML, download in bulk, or query the metadata as RDF (facts written as subject, predicate, and object statements) through a SPARQL endpoint and a dedicated web service. Anchor everything on the CELEX number: it is the stable key that lets you fetch a document, join it to related acts, and spot duplicates when the same text arrives through more than one channel.
The headline strength for retrieval is that the identical legal text exists in all 24 official languages. You can answer a Lithuanian or Greek query from the authoritative wording, or align the same article across languages so a cross-lingual retriever returns matching passages. That makes it a natural base for compliance assistants, cross-border research tools, and anything that must quote the real law rather than paraphrase it.
Plan for a few rough edges. A consolidated version is convenient but is editorial, not always the legally binding text, so store the CELEX identifier and the version date with every chunk. Markup quality improves over time, and older scanned acts are messier than recent XML. Licensing is refreshingly clear: reuse is authorised under Commission Decision 2011/833/EU for commercial and non-commercial work, with no share-alike obligation, though a source credit is courteous.
Treat EUR-Lex as your canonical EU layer and add Legislation.gov.uk for UK statute or national gazettes where you need member-state depth.
Related sources
AustLII
A free resource of medium-neutral case law and unreported judgments for all Australian jurisdictions, covering the whole country since 1995.
Awesome LegalTech
A curated list of legal technology resources: open-source platforms, AI models, companies, datasets, and tools spanning the global legal ecosystem. Useful for tracking new entrants in a fast-moving space.
BAILII (British and Irish Legal Information Institute)
Free access to British and Irish primary legal materials, covering UK and Ireland case law and legislation. Alongside the National Archives, one of the main free sources for reading UK judgments.
Cambridge Law Corpus
A research dataset of more than 250,000 UK court cases, mostly from the 21st century but with some reaching back to the 16th century. Built by the University of Cambridge for legal natural language processing work, it is available under restricted access for research use.