CORE
CORE pulls together open-access research from institutional repositories, preprint servers, and journals into a single searchable collection. Instead of querying each repository separately, you get one place to find and retrieve full-text papers across disciplines.
It exposes the collection several ways: a search API for live lookups, bulk data dumps for offline processing, and a dataset service for research use. That mix suits both quick prototypes and full corpus builds.
Licensing depends on each paper, since CORE aggregates rather than publishes, so full-text reuse terms follow the original source. The metadata itself is freely available. For RAG, CORE is a practical way to gather broad open-access coverage without integrating dozens of repositories by hand.
Related sources
arXiv
An open-access preprint server for physics, mathematics, computer science, quantitative biology, statistics, and more, holding over 2.5 million papers. You can pull the full archive in bulk from Amazon S3, or harvest the metadata through OAI-PMH, a standard protocol for sharing records between repositories.
Crossref
A nonprofit DOI registration agency that publishes metadata for over 150 million scholarly works, including journal articles, books, and conference proceedings. A DOI is the permanent identifier assigned to each work, and Crossref shares its metadata through a free API.
OpenAlex
A free, open catalogue of the world's scholarly works, authors, venues, institutions, and research topics. It succeeds Microsoft Academic Graph and holds hundreds of millions of works with rich metadata and citation links.
PubMed / PubMed Central (PMC)
The US National Library of Medicine's database of biomedical and life sciences literature. PubMed indexes over 36 million citations and abstracts, and PubMed Central (PMC) adds free full-text access to a growing subset. You can download the data in bulk over FTP.