Skip to content
RAG Repo

Harvard USPTO Patent Dataset (HUPD)

The Harvard USPTO Patent Dataset (HUPD) is a corpus of US patent applications put together for machine learning and natural language processing (NLP) work, the field of teaching computers to work with human language.

Most popular patent search tools are built for lawyers and analysts rather than for researchers training models, so their data is awkward to work with at scale. HUPD addresses that by providing a well-structured, ready-to-use collection that slots neatly into ML and NLP pipelines.

It is hosted on Hugging Face under CC BY 4.0, so you can load it directly with standard dataset tooling and use it commercially as long as you credit the source.

patentsususptonlpmachine-learningresearch

Related sources