Skip to content
RAG Repo

NLP Datasets is a focused GitHub list that catalogues datasets made of text, the raw material for most retrieval and language work. Entries span general corpora, conversational and dialogue data, sentiment-labelled sets, and summarisation collections, each with a short note and a link to the source.

The alphabetical layout makes it easy to skim, though it means you browse by name rather than by task. As with any community list, check each dataset's licence and last update at its home before you rely on it. It pairs well with broader directories when you specifically need language data rather than tabular or geospatial sets.

awesome-listnlptext-corporadirectorygithub

Related sources