SPICE 1.1.2: Quantum Mechanical Data for Drug-like Molecules and Peptides
by Peter Eastman / Stanford University
Available on 1 platform
Sign in to view source links and access this dataset
Description
SPICE is a collection of quantum mechanical data for training machine learning potential functions, with an emphasis on simulating drug-like small molecules interacting with proteins. It was created by researchers including Peter Eastman from Stanford University and described in a 2022 publication. The data includes molecular conformations, energies, gradients, and multipole moments for molecules identified by PubChem IDs, SMILES strings, or amino acid sequences.
Use Cases
Training machine learning potential functions based on quantum mechanical energy and gradient data.
Simulating interactions between drug-like molecules and proteins using the provided conformational and atomic property data.
Analyzing molecular charge and multipole distributions using the MBIS charge, dipole, quadrupole, and octupole arrays.
Strengths
Data is structured for machine learning with specific arrays for conformations, energies, and gradients.
Includes multiple quantum mechanical properties per conformation, such as formation energy, DFT total energy, and molecular multipoles.
Molecules are sourced from specific subsets like PubChem and include canonical SMILES strings with explicit hydrogens.
Limitations
Row count and total dataset size are unknown, which may limit suitability assessment.
Some data groups may be missing certain arrays if the MBIS calculation failed to converge.
Last update date is unknown; freshness unverified.
Provenance
Source
Stanford University researchers, as described in the 2022 arXiv publication.
Collection Method
Quantum mechanical calculations, likely Density Functional Theory (DFT), for generated molecular conformations.
Data is stored in an HDF5 file with a specific hierarchical structure; users need compatible libraries to read it. The license is listed as Open Access (green).