Cultural heritage collections are the digitised holdings of museums, libraries, and archives: scanned books and manuscripts, photographs of artefacts, catalogue records, and the descriptive notes that curators have written over decades. For a RAG system, a setup that answers questions by first retrieving relevant passages and then generating a reply from them, this is unusually clean raw material. The metadata (the structured facts about each item, like its date, creator, and place of origin) is often richer and more carefully checked than anything you will find on the open web, which makes retrieval more precise.
Choosing within this category comes down to what you are actually retrieving. Aggregators like Europeana and the Digital Public Library of America pull records from thousands of institutions, so you get breadth but variable depth. A single institution such as the Smithsonian gives you consistency and fuller descriptions across a narrower collection. Decide whether you want the metadata, the full text, or the images, because many of these sources are strong on one and thin on the others.
The licensing needs care. βPublic domainβ usually refers to the underlying work, an old painting or a nineteenth-century book, while the photograph or scan of it may carry its own rights, and the catalogue text almost always does. Check the rights statement on the item, not just the collectionβs headline, and watch for mixed collections where a handful of items are still in copyright. The sources below range from broad cross-institution aggregators to single-museum archives with deep, well-described holdings.