Loading...
Loading...
General ML benchmarks, tabular data, AutoML, recommendation systems, anomaly detection, evaluation suites
193,236 datasets
2534 deep convection initiation (DCI) events identified in the North China area during the summer seasons (June, July, and August) from 2017 to 2019. The dataset was collected by an automatic DCI identification method using Himawari-8 satellite infrared data at 5 km resolution and 10-minute intervals. It was authored by G. Lu of Peking University and sourced from the paperswithcode platform.
A neural network-based tool for predicting formation energies of materials based on elemental and structural features. The models were trained on data from the Open Quantum Materials Database (OQMD) and achieve mean average errors between 28 and 42 meV/atom on a test set of 21,800 compounds. The work was authored by Adam M. Krajewski of Pennsylvania State University.
Ten datasets contain logs, traces, and KPI data from the Train-Ticket microservice benchmark system. The data was collected by Monika Steidl at Universität Innsbruck and includes explanations of identified anomalies. Each dataset folder is structured by the changed microservice, third-party library, version, and collection date.
A field experiment tested a Beta-Delta model of present bias in saving decisions among low-income tax filers. The study, conducted by Damon Jones of UC Berkeley, found qualitative evidence of present-biased preferences, estimating parameters of 0.34 and 1.08 over an 8-month horizon. This translates to an annual discount rate of 164%.
268 survey responses evaluate Continuing Medical Education programs from 2021–2022. The data includes participant demographics, overall satisfaction scores, and criterion-based evaluations across course phases. Hela Ghali published this dataset on figshare under a CC-BY-4.0 license.
Sergej Fries of RWTH Aachen University proposes modifications to the P3C projected clustering algorithm for large, high-dimensional data sets. The work includes the novel P3C+-MR algorithm and its simplified P3C+-MR-Light variant, both implemented in MapReduce. Their effectiveness and efficiency are evaluated on synthetic and real-world data sets.
ScriptNet's cBAD dataset contains the training and test set for the ICDAR 2017 Competition on Baseline Detection in Archival Documents. It comprises 2035 annotated document page images collected from 9 different archives, split into two tracks: one for simple handwritten paragraphs and another for complex documents with tables, marginalia, and noise. The training data includes PAGE XML files with manually annotated text regions and baselines.
WRF-SUEWS model input data supports evaluation at KCL and SWD sites in the UK. The archive is associated with a model development paper and its code is available on GitHub. Ting Sun from University College London authored this dataset.
A simulated dataset of silicon pixel clusters produced by charged pions, where particle kinematics are taken from fitted tracks in CMS Run 2 data. The dataset includes truth properties for each particle and 3D charge deposition clusters across time slices. It was created by M. Swartz of Johns Hopkins University.
Records from approximately 130,000 to 75,000 years ago capture sea level using U-series dated cave deposits. The data is a global standardized database exported from The World Atlas of Last Interglacial Shorelines (WALIS). It was authored by Oana-Alexandra Dumitru of Columbia University.
A global standardized database of U-series dated speleothems capturing sea level during the last interglacial period. It was exported from The World Atlas of Last Interglacial Shorelines (WALIS) database. The dataset was authored by Oana-Alexandra Dumitru of Columbia University.
Two large-scale virtual screening datasets for benchmarking machine learning methods in early-phase drug discovery. The datasets, provided by Andreas Luttens of Uppsala University, contain canonical SMILES, compound identifiers, and docking scores for approximately 15.5 million 'Rule-of-Four' molecules and approximately 235 million 'lead-like' molecules docked against eight different biological targets.
Gilbert J. Botvin of Cornell University authored a paper discussing the definition and measurement of success in drug abuse prevention. The paper argues for viewing success as a process of incremental progress and suggests benchmarks for assessment. It notes that considerable, though slow, progress has been made over the past decade.
A Gaussian approximation potential (GAP) for modeling amorphous carbon structures. The potential was fitted using the QUIP/GAP framework by recomputing the a-C database of Deringer and Csányi at the PBE+MBD level of theory with the VASP code. It employs 2-body, 3-body, and SOAP-type descriptors from the TurboGAP code and is compatible with both QUIP/GAP and TurboGAP software.
A Gaussian approximation potential for silicon fitted using the QUIP/GAP and TurboGAP codes. The potential uses 2-body, 3-body, and SOAP-type descriptors, based on a recomputed database from Bartók et al. at the PW91 level of theory. It was created by A. Miguel of Aalto University.
A dataset from Leicester City Council's Local Plan 2022 consultation evaluating neighbourhood parades. It likely contains scores based on facility points (convenience store, Post Office, pharmacy), ATM presence, percentage of national operators, and vacancy rates. The data was last updated on 2026-06-17.
Charnwood Borough's official 2020 register catalogs previously developed land available for redevelopment. This dataset likely contains site-specific details useful for assessing development potential and planning policy compliance. Its publication under an open license facilitates analysis by planners, developers, and researchers.
Experimental data and simulation scripts reproduce the results of the paper 'Bloch point-mediated skyrmion annihilation in three dimensions' by Birch et al. The dataset, published by David Cortés‐Ortuño of Utrecht University in 2021, is hosted on Zenodo and linked to a GitHub repository for updates. It supports research into three-dimensional magnetic texture dynamics.
1,517,419 quantum reaction rate constant products computed from transmission coefficients for model single and double barrier minimum energy paths. The dataset was created by Evan Komp and Stacey Valleau at the University of Washington for a 2020 publication to train a deep neural network. It includes features like barrier widths, heights, symmetry constants, and temperature.
Data associated with the GitHub repository for reproducing analysis and figures from Sethi et al. (in prep.). The dataset is authored by Sarab S. Sethi of Imperial College London and was released in November 2019. It is provided under an Open Access license.