Open Molecules 2025 (OMol25)
Open Molecules 2025 (OMol25) is a dataset released by Meta's FAIR Chemistry team, containing over 100 million single-point density functional theory (DFT) calculations computed at the wB97M-V/def2-TZVPD level of theory using ORCA6. Each structure is labelled with total energy (in eV) and atomic forces (in eV per angstrom). The dataset blends elemental, chemical and structural diversity: it covers 83 elements, organic and inorganic molecules, transition metal complexes and electrolytes, explicit solvation, variable charge and spin, conformers and reactive structures, with systems of up to 350 atoms. It represents billions of CPU core-hours of compute.
Access is through HuggingFace, where the dataset is gated behind a request form and provided as ASE DB compatible LMDB files (.aselmdb), with charge and spin multiplicity stored in the atoms.info dictionary. The accompanying fairchem library and documentation cover loading and training workflows, and Meta also released a family of eSEN models trained on the data.
This is a specialist scientific resource rather than a text corpus. It is best suited to training or evaluating machine-learning interatomic potentials and foundation models for chemistry and materials, and for supplying structured facts to a domain RAG system through a metadata layer or knowledge graph rather than raw retrieval. Watch-outs: the dataset licence (CC BY 4.0) requires attribution, access is gated and geographically restricted (unavailable in sanctioned jurisdictions and in China, Russia and Belarus), and the released model checkpoints carry a separate proprietary FAIR Chemistry Licence, so do not assume the models share the data terms. Related sources in this catalogue include Open Materials 2024 (OMat24) for inorganic materials and the broader scientific holdings of the Materials Project.
Related sources
Bio2RDF
An open-source project that pulls together a diverse set of life-sciences datasets from many providers into a single linked-data graph, with a SPARQL endpoint for querying across them. The full collection is about 11 billion triples across 35 datasets, including DrugBank, PubMed, and MeSH.
ChEBI
A dictionary of small chemical compound molecular entities, with structure files and ontology files available for download. Particularly useful as a controlled vocabulary for chemical entity linking, where you match a mention in text to a standard identifier.
ChEMBL
A manually curated database of bioactive molecules with drug-like properties, bringing together chemical, bioactivity, and genomic data to support drug discovery. Holds close to 2.5M compound records on nearly 2M unique chemical structures, extracted mainly from the medicinal chemistry literature.
Crystallography Open Database
A collection of over 350,000 crystal structure files covering organic, inorganic, and metal-organic compounds, released into the public domain under CC0.