Skip to content
RAG Repo

PD12M (Public Domain 12M)

PD12M is the largest image-text dataset drawn exclusively from public domain and CC0 material, which means the source images carry no copyright restrictions. At 12.4M image-caption pairs it matches the scale of widely used web-scraped datasets, so you can train comparable models without the licensing risk that comes with scraping images under unknown terms.

The images are hosted on dedicated cloud storage separate from their original websites. This addresses link rot, the slow decay of web-scraped datasets as the pages they point to go offline, which leaves many older image-text datasets full of dead links over time.

A smaller 3.3M item subset, PD3M, is available for teams that want the same clean licensing at a more manageable size.

image-textpublic-domaincc0multimodalcommercial-friendly

Related sources