GDELT
The Global Database of Events, Language, and Tone (GDELT) is an ongoing catalogue of what the world's news is reporting. It parses articles in real time, pulls out the events, people, organisations, locations, and themes they mention, and scores their emotional tone. New data is added every 15 minutes.
Two datasets sit at its core. The Event Database records who did what to whom, where, and when, using a standard coding scheme, stretching back to 1979. The Global Knowledge Graph (GKG) connects the entities, themes, and emotions found across coverage, which lets you trace how a story or topic moves through the media over time.
You can query GDELT directly on Google BigQuery, download raw files, or use its APIs, with everything published as CSV. For RAG, it is a rich source for temporal and current-affairs questions, media analysis, and tracking how events and sentiment evolve across languages and countries.
Related sources
CC-News
A subset of Common Crawl focused on news, containing millions of articles pulled from news websites around the world. It has fed several large language model training pipelines and gives you a ready news corpus without crawling sites yourself.
Internet Archive
A nonprofit digital library giving free access to millions of books, films, audio recordings, software, archived websites, and television news. Its Wayback Machine has saved more than 800 billion web pages, making it a deep well of historical and current text.