Loading...
Loading...
Organic/inorganic chemistry, analytical chemistry, electrochemistry, molecular properties, chemical reactions
2,293 datasets
A study from the Australian Ocean Data Network combines comprehensive two-dimensional gas chromatography (GCxGC-TOFMS) and compound-specific isotope analyses (CSIA) of diamondoids to identify hydrocarbon contributions in the Browse Basin. The analysis includes non-biodegraded (Caswell) and biodegraded (Cornea and Gwydion) oil fields, using diamondoid concentrations and isotopic composition as source-specific indicators resistant to thermal maturity and biodegradation. The dataset, last updated in 2026, is provided in PDF format.
An algorithm-generated dataset from the University of Bonn creates large compound profiling matrices. These matrices record assay results for compound libraries tested against panels of targets, derived from the PubChem BioAssay database. The methodology can generate matrices containing more than 100,000 compounds exhaustively tested against 50 to 100 targets.
A collection of 8 lists of chemicals associated with Parkinson's disease and related disorders, created by Begoña Talavera Andújar at the University of Luxembourg. The lists were generated by exploring PubChem functionality that links Medical Subject Headings (MeSH) disease information to chemicals. Details are available in a published manuscript (https://doi.org/10.1007/s00216-022-04207-z).
A manually validated list of metabolites and transformation products extracted from the Hazardous Substance Data Bank (HSDB) in PubChem. The dataset is actively maintained, with updates from 2020 through 2023 adding new substances, reactions, and columns like Biosystem and Enzyme. It serves as a curated resource for suspect screening and environmental fate studies of hazardous chemicals.
Input and output files for Ab Initio Molecular Dynamics simulations of Li6PS5X (X=I, Cl) argyrodite materials, as described in the associated research article. The dataset includes VASP input files and outputs for full AIMD runs and for derived inherent structure trajectory calculations. It was created by Benjamin J. Morgan of the University of Bath.
A collaborative study evaluated crude protein content in four feed samples across seven laboratories. Each laboratory analyzed samples over six days with three replicates per sample per day. The study found significant variation among laboratories, which accounted for 43.6% to 68.2% of the total random variation.
Result files from the Aalto University study 'Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data' by Eric Bach. The dataset contains raw and averaged max-marginal predictions, top-k accuracies, and rank analyses for the LC-MS²Struct method and comparison methods. It includes two versions of results based on 2D and 3D molecular fingerprints.
Data generated by multiple DFT programs to optimize the structure of sodium peroxodisulfate and calculate its infrared, attenuated total reflectance, and terahertz absorption spectra. The dataset, authored by John Kendrick of the University of Leeds, compares results from different computational methods and effective medium models against experimental measurements. It highlights the sensitivity of low-frequency phonon modes to calculation parameters like k-point grids and energy cutoffs.
1,215 compounds from the ChEMBL database were used to construct multi-target QSAR models for predicting inhibitory activity against class I HDAC isoforms. The models, including linear and non-linear classifiers, were built using 13 deviation descriptors and achieved accuracies exceeding 90% on sub-training, test, and validation sets. Author G.G. Tu published the dataset on figshare in May 2026, which includes results from virtual screening, drug-likeness assessment, and molecular dynamics simulations.
Bus stop data for Bristol includes Naptan codes, names, locations, and amenity indicators for seating, shelter type, and raised kerbs. The dataset is provided by Bristol City Council and was last updated on 2026-07-08. It references the UK-wide Naptan code system maintained by the Department for Transport.
Annotated compound data sets are part of a freely available database and software tool collection developed at the University of Bonn. The original release was described in a 2012 F1000Research article, with an updated version described in a forthcoming data note. The data is provided by author Ye Hu.
A freely available database of annotated compound data sets and software tools for chemoinformatics and computational medicinal chemistry. The programs were developed in a laboratory at the University of Bonn by Ye Hu. The original release was described in a 2012 article in F1000Research.
Real-world vehicle emissions of carbonyl compounds were measured in the Rebouças Tunnel in Rio de Janeiro, Brazil. The dataset includes 16 samples collected at two points inside the 2840-meter tunnel on 8 different days, with approximately 5,000 vehicles per hour passing through. Emission factors and ozone-forming potentials for formaldehyde and acetaldehyde were calculated using two methods, with results reported by José Claudino Souza Almeida.
Browse Basin, North West Shelf, Australia, hydrocarbon fluid samples analyzed using comprehensive two-dimensional gas chromatography coupled to time-of-flight mass spectrometry (GC × GC-TOFMS). The dataset includes diamondoid and semi-volatile aromatic compound data, aromatic maturity ratios, and stable carbon isotopic compositions to identify mixed hydrocarbon sources. The study, published in Marine and Petroleum Geology in 2020, reveals a complex fill history for accumulations including the greater Cornea field and Gwydion-1.
91 real milk samples were analyzed for three antibiotics using an optimized QuEChERS extraction method and ultra-performance liquid chromatography-tandem mass spectrometry. The method achieved recoveries from 95 to 99%, linearity (R2) above 0.96, and limits of detection between 1.4-6.8 µg L-1. This research dataset was contributed by Andressa Grabsk via the paperswithcode platform.
Regular updates of the PubChemLite for Exposomics data collection, a subset of PubChem compiled from 11 major categories. The repository, maintained by Emma Schymanski at the University of Luxembourg, provides a curated list of PubChem Compound Identifiers collapsed by InChIKey first block and filtered for compatibility with metabolite identification tools. The collection is described in a 2021 publication (DOI:10.1186/s13321-021-00489-0).
A method for determining L-ascorbic acid in milk using UHPLC-MS/MS was validated for linearity, accuracy, and precision. The study reports a limit of detection of 1.5 µg L⁻¹ and a limit of quantification of 5.0 µg L⁻¹. The method was applied to milk samples by author Caroline Zappielo.
A research paper describes an optimized solid-phase microextraction method using cork as a biosorbent for detecting two UV filters, 4-MBC and OD-PABA, in river water. The method achieved quantification limits of 0.1 and 0.01 µg L-1, respectively, with recovery values between 67 and 107%. The work was authored by Ana C. Silva and published on the PaperswithCode platform.
An extensive assessment of DFT+U band gap predictions for 20 compounds containing transition-metal or p-block elements. The study compares results using nonorthogonalized and orthogonalized atomic orbitals as Hubbard projectors, finding the latter provide the most accurate gaps. This work by Nicole E. Kirchner-Hall demonstrates a method for reliable band gap prediction at moderate computational cost.
34 species of fire emissions data from the Fire Modeling Intercomparison Project (FireMIP) covering the years 1700 to 2012. The dataset contains multi-model estimates of global fire emissions for elements, compounds, and classes of compounds. It was authored by Fang Li of the Chinese Academy of Sciences and published in 2019.