Systematic Analysis of Single- and Multi-Target Compounds for Machine Learning
by Christian Feldmann / University of Bonn
Available on 1 platform
Sign in to view source links and access this dataset
Description
Two balanced datasets contain 15,142 multi-target and 15,081 single-target compounds, plus a subset of 1,828 diverse-target and 1,776 single-target compounds. The data was compiled by Christian Feldmann at the University of Bonn for a study published in Molecular Pharmaceutics. Each compound includes a SMILES representation, ChEMBL ID, UniProt target IDs, and a category label.
Use Cases
Train classification models to distinguish single-target from multi-target compounds based on SMILES representations.
Analyze structural patterns associated with diverse-target activity using the provided compound subset.
Benchmark feature selection or data reduction techniques using the 'Y' tag indicating compound retention after random or nearest-neighbor removal.
Strengths
Contains over 30,000 total compounds across two balanced datasets.