5,683 records combine data from PubChem and ChEMBL, two major public chemistry databases. The dataset likely contains information on proteins and their associated small molecules. It was published on Kaggle, but the original author and specific collection details are not provided.
Use Cases
- Train a model to predict protein-ligand binding affinity (inferred from domain, verify after download)
- Screen for potential drug candidates by analyzing molecular properties (inferred from domain, verify after download)
- Build a knowledge graph linking proteins, compounds, and bioactivities (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform with established data sharing infrastructure.
- Explicitly states a count of 5,683 records.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Last update date is unknown; freshness unverified.
Provenance
- Source
- Aggregated from PubChem and ChEMBL public databases.