Loading...
Loading...
Organic/inorganic chemistry, analytical chemistry, electrochemistry, molecular properties, chemical reactions
2,170 datasets
Canadian Food Inspection Agency data from a 5-year targeted survey of dioxins and dioxin-like compounds in selected foods. Between April 2014 and March 2019, 3,115 samples were collected from retail stores across Canada, with compounds detected in 95% of samples. Health Canada's Bureau of Chemical Safety evaluated the levels and determined none posed a risk to human health.
SPICE is a collection of quantum mechanical data for training machine learning potential functions, with an emphasis on simulating drug-like small molecules interacting with proteins. It was created by researchers including Peter Eastman from Stanford University and described in a 2022 publication. The data includes molecular conformations, energies, gradients, and multipole moments for molecules identified by PubChem IDs, SMILES strings, or amino acid sequences.
A single-sensor approach identifies volatile organic compounds by analyzing non-equilibrium mass-transport dynamics. Machine learning models trained on the sensor's time-dependent reflectance spectra can predict chemical properties like boiling point and vapor pressure. The data originates from research at Harvard University, implementing both passive and active sniffing modalities.
Issar Arab and Khaled Barakat's dataset contains 8879 unique molecular compounds for predicting hERG potassium channel cardiotoxicity liability. The data, gathered from ChEMBL, PubChem, and literature, provides SMILES strings and corresponding pIC50 potency values. It is pre-split into 8380 training and 499 testing compounds for building descriptor-based machine learning models.
A diverse reaction database (RDB7) with transition states containing up to 7 heavy atoms, published in 2022. The dataset contains atom-mapped SMILES, barrier heights, reaction enthalpies, and rate coefficients for thousands of reactions, calculated using quantum chemistry methods. It was created by researchers at the Massachusetts Institute of Technology and includes raw computational output files.
This dataset contains computational quantum chemistry data for elementary chemical reactions. It includes Q-Chem output files, atom-mapped SMILES strings, activation energies, and enthalpies of formation for 16,452 reactions calculated with the B97-D3/def2-mSVP method and for 12,001 reactions calculated with the ωB97X-D3/def2-TZVP method. The raw data comprises geometry optimization and harmonic vibrational analysis log files for reactants, products, and transition states, with separate archives containing 69,366 B97-D3/def2-mSVP and 24,987 ωB97X-D3/def2-TZVP transition state calculations.
Nicolas Mounet and École Polytechnique Fédérale de Lausanne created a database of 2D materials identified through computational screening. Starting from 108,423 experimentally known 3D compounds, the work identified 1,825 potentially exfoliable materials, including 1,036 easily exfoliable cases. For 258 compounds, the database includes calculated properties such as vibrational, electronic, magnetic, and topological data.
Imperial College London research data underpins the article on [2.2.2.2]Paracyclophanetetraenes (PCTs). The dataset contains experimental spectroscopic and electrochemical measurements for 11 molecules, including Br-P2, O-PCT, and S-PCT, alongside computational chemistry files for two molecules. It provides raw measurement files and quantum chemistry input/output files for geometry optimizations of various electronic states and rotamers.
Experimental data from angle-resolved photoemission spectroscopy and low-energy electron diffraction measurements on a cuprous oxide film. The dataset includes raw and processed spectra files from proprietary and open-source analysis software, such as Igor Pro and SpecsLab Prodigy. Jan Beckord from the University of Zurich authored the associated research paper.
PubChemLite for Exposomics is a curated subset of PubChem, compiled from 11 annotation categories including AgroChemInfo, ToxicityInfo, and FoodRelated. It includes predicted collision cross section values for 11 adducts, calculated using the c3sdb code from CCSbase. The dataset is maintained by Emma Schymanski at the University of Luxembourg and is updated monthly, with the latest version available via a Zenodo DOI.
A training dataset for the OrbNet Denali machine learning potential, consisting of molecular geometries and corresponding energy labels. The data includes geometries in XYZ+ format and energy labels calculated at the wB97X-D3/def2-TZVP and GFN1-xTB levels of theory. It was created by researchers from Entos, Inc., Caltech, and NVIDIA for the 2021 paper "OrbNet Denali: A machine learning potential for biological and organic chemistry with semi-empirical cost and DFT accuracy".
A University of Oxford study by V. Fulop details the structural and spectroscopic changes in cytochrome P450cam from Pseudomonas putida after modification with an electroactive reagent. The dataset likely contains results from X-ray crystallography at 2.2 Å resolution and electrochemistry experiments. It describes the attachment of two ferrocene molecules, conformational changes in the enzyme's active site, and a shift in the Soret absorption band.
Oxygen plasma post-treatment enhanced the photo-corrosion resistance of ZnO nanowire films and led to a 46% higher degradation of phenol compared to as-produced films. This dataset contains photodegradation, phenol calibration, Zn concentration, X-ray diffraction, hydrodynamics calculations, and TEM/SEM images supporting the results. The data was produced by Caitlin A. Taylor at the University of Bath.
The Browse Basin offshore Northwest Australia is a major hydrocarbon province. This dataset likely contains results from comprehensive two-dimensional gas chromatography and compound-specific isotope analyses of diamondoids and other hydrocarbons, used to unravel complex charge histories. The data was published by the Australian Ocean Data Network and last updated in 2026.
Two balanced datasets contain 15,142 multi-target and 15,081 single-target compounds, plus a subset of 1,828 diverse-target and 1,776 single-target compounds. The data was compiled by Christian Feldmann at the University of Bonn for a study published in Molecular Pharmaceutics. Each compound includes a SMILES representation, ChEMBL ID, UniProt target IDs, and a category label.
A collection of promiscuity cliffs (PCs), promiscuity cliff pathways (PCPs), and promiscuity hubs (PHs) formed by kinase inhibitors, covering more than 80% of the human kinome. The data structures were introduced for analyzing compound promiscuity and are provided by Filip Miljković from the University of Bonn. The dataset is referenced in multiple published research articles.
PubChemLite for Exposomics is a curated subset of PubChem compiled from 10 annotation categories relevant to environmental exposure science. The dataset, created by Emma Schymanski at the University of Luxembourg, includes predicted collision cross section (CCS) values for 8 adducts, provided by the CCSbase project. This version was released on 1 January 2021, with a newer version available for monthly download.
10,000 algorithmically generated CsPb(Cl/Br)3 perovskite alloy structures with single-point DFT calculations across four space groups. The data was created by Jarno Laakso of Aalto University to train a machine learning model for predicting material stability. It includes four subsets covering single-point energies, extended concentration tails, relaxation snapshots, and pure endpoint structures.
Gas chromatography-mass spectrometry (GC × GC-TOFMS) analysis of semi-volatile aromatics and diamondoids from whole oil/condensate samples reveals multiple hydrocarbon sources in Browse Basin accumulations. The data identifies high-maturity fluids in biodegraded accumulations on the Yampi Shelf and suggests complex fill histories involving Jurassic source rocks like the Plover and Vulcan Formations. Results provide supplementary information to traditional saturated biomarker-based fluid discrimination.
A study from the Australian Ocean Data Network combines comprehensive two-dimensional gas chromatography (GCxGC-TOFMS) and compound-specific isotope analyses (CSIA) of diamondoids to identify hydrocarbon contributions in the Browse Basin. The analysis includes non-biodegraded (Caswell) and biodegraded (Cornea and Gwydion) oil fields, using diamondoid concentrations and isotopic composition as source-specific indicators resistant to thermal maturity and biodegradation. The dataset, last updated in 2026, is provided in PDF format.