ChEMBL
ChEMBL connects chemical structures, their measured biological activity, and genomic information, helping researchers translate genomic findings into new drugs. New releases come out roughly every three to four months.
Unlike bulk databases, ChEMBL data is abstracted from the literature and curated by hand. That makes it smaller than automatically gathered collections but higher in confidence, which matters when the answers feed clinical or medicinal chemistry work.
ChEMBL is published under CC BY-SA 3.0. You can use it commercially and you must credit the source, but the share-alike term means any derivative database you distribute has to carry the same licence. Whether embeddings (numeric vectors derived from the text) count as a derivative under share-alike is genuinely unsettled, so factor that into your plans.
Related sources
Bio2RDF
An open-source project that pulls together a diverse set of life-sciences datasets from many providers into a single linked-data graph, with a SPARQL endpoint for querying across them. The full collection is about 11 billion triples across 35 datasets, including DrugBank, PubMed, and MeSH.
ChEBI
A dictionary of small chemical compound molecular entities, with structure files and ontology files available for download. Particularly useful as a controlled vocabulary for chemical entity linking, where you match a mention in text to a standard identifier.
Crystallography Open Database
A collection of over 350,000 crystal structure files covering organic, inorganic, and metal-organic compounds, released into the public domain under CC0.
Materials Project
A database of computed properties for known and predicted materials, produced through high-throughput density functional theory calculations (a physics method for modelling how electrons behave in a material). Free API access with registration.