Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Comprising approximately 36,000 unique pairs of protein sequences and ligand SMILES strings, along with the 3D coordinates of their complexes from the Protein Data Bank (PDB). Ligands are filtered to have at least 3 atoms, a molecular weight of 100 Da or more, and exclude the 280 most common PDB ligands. It was created by author jglaser and last updated in October 2022.
The full description and detailed metadata are available on the Hugging Face dataset page. The ligand SMILES strings are assumed to be tokenized using the regex pattern from P. Schwaller.