Sign in to view source links and access this dataset
Description
22,153 AI-generated, novel, drug-like small molecules are provided as RDKit-verified SMILES strings. The dataset was created by MKEChem and was last updated on July 6, 2026. Each molecule has been validated for chemical validity and novelty against 4,643,595 known compounds from sources like MOSES, ZINC-250k, and ChEMBL.
Use Cases
Training generative AI models for novel molecule design based on the AI-generated SMILES strings.
Benchmarking virtual screening pipelines against a set of pre-filtered, drug-like compounds.
Exploring chemical space for novel scaffolds in early-stage drug discovery based on the validated novelty claim.
Strengths
Contains 22,153 molecules, providing a substantial set for model training or analysis.
Novelty is verified against a large reference set of 4,643,595 known compounds from established databases.
All molecules are described as chemically valid and drug-like, indicating pre-filtering for pharmaceutical relevance.
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
The dataset is listed as a commercial product for purchase, which may restrict access.
Provenance
Source
MKEChem, via Hugging Face.
Collection Method
AI-generated, with subsequent validation and filtering for chemical validity and novelty.
Freshness
Last updated 2026-07-06 13:41:25; freshness should be verified.
License is unknown. The dataset is listed as a commercial product available for purchase on Gumroad, not freely downloadable.