PubChem
PubChem is one of the largest open chemistry resources in the world. Most of its records are small molecules, but it also holds larger molecules such as nucleotides, carbohydrates, lipids, peptides, and chemically modified macromolecules. For each entry you get chemical structures, standard identifiers, measured physical properties, biological activities, and links to patents and safety data.
For structured retrieval, the PubChemRDF project publishes the Compound, Substance, and Bioassay databases as RDF, a way of expressing data as subject-predicate-object statements called triples. That amounts to roughly 80 billion triples, which you can load into a knowledge graph or query directly.
As a work of the US Government, PubChem is in the public domain, so you can use it freely in commercial products with no licence conditions attached.
Related sources
Bio2RDF
An open-source project that pulls together a diverse set of life-sciences datasets from many providers into a single linked-data graph, with a SPARQL endpoint for querying across them. The full collection is about 11 billion triples across 35 datasets, including DrugBank, PubMed, and MeSH.
ChEBI
A dictionary of small chemical compound molecular entities, with structure files and ontology files available for download. Particularly useful as a controlled vocabulary for chemical entity linking, where you match a mention in text to a standard identifier.
ChEMBL
A manually curated database of bioactive molecules with drug-like properties, bringing together chemical, bioactivity, and genomic data to support drug discovery. Holds close to 2.5M compound records on nearly 2M unique chemical structures, extracted mainly from the medicinal chemistry literature.
Crystallography Open Database
A collection of over 350,000 crystal structure files covering organic, inorganic, and metal-organic compounds, released into the public domain under CC0.