HathiTrust Digital Library
HathiTrust is a shared digital library run by a large partnership of academic and research institutions, holding millions of digitised books, journals, and other titles pooled from member libraries' scanning programmes. Public domain works can be read and often downloaded in full; in-copyright material is preserved and searchable but not freely readable.
The route that matters for data work is the HathiTrust Research Center (HTRC), which opens the whole collection, in-copyright titles included, for computational analysis under a "non-consumptive" model. Non-consumptive means you can run analysis across the texts, counting, modelling, and extracting features, without being able to read or lift out the original passages yourself. The Research Center supplies the datasets and a secure environment for this work, along with pre-computed 'Extracted Features', per-page word counts and metadata you can download openly for millions of volumes.
This shapes what HathiTrust is genuinely good for. It is a research resource for studying patterns across an enormous book corpus, not a well of reusable passages you can drop into a vector database. If your project needs to retrieve and quote text, only the public domain portion is usable that way, and even then you will be assembling it title by title.
Licensing is the big caveat, so read it carefully. Terms vary by title, in-copyright content is available only under non-consumptive research terms that rule out commercial reuse, and the Extracted Features datasets, though openly licensed, are deliberately built so you cannot reconstruct the original text. Treat commercial use as restricted unless you have confirmed that a specific title is public domain.
For freely reusable book text you can retrieve and quote, Project Gutenberg and the Internet Archive are the more practical starting points. Come to HathiTrust when the scale and scholarly breadth of its corpus, and its computational-access model, are exactly what your research needs.
Related sources
Common European Data Space for Cultural Heritage
The umbrella project coordinating cultural heritage data across Europe. Led by the Europeana Initiative, it helps cultural institutions and EU Member States share their collections through a common, shared infrastructure.
Digital Public Library of America (DPLA)
An aggregator that brings together content from America's museums, libraries, and archives. Its API offers metadata on individual items and on collections, and the whole repository is available to download as zipped JSON files.
Europeana
A platform for discovering cultural heritage collections across Europe, with multilingual access to over 60 million digitised items from around 4,000 institutions, including books, paintings, maps, manuscripts, and audiovisual and 3D media.
Smithsonian Open Access
Millions of digital items from the Smithsonian's museums, research centres, and archives, released under CC0. Includes images, 3D models, and research datasets you can reuse freely, including commercially.