Loading...
Loading...
General ML benchmarks, tabular data, AutoML, recommendation systems, anomaly detection, evaluation suites
194,257 datasets
A 2026-07-08 validated fork of the Scale-SWE dataset, containing 17,202 Python issue-resolving tasks that produce a clean reward signal end-to-end. The dataset was created by PrimeIntellect by removing 2,979 rows (14.8%) from the upstream version through validation against exclusion lists and failure categorization.
Yibo Li published a research dataset on figshare in June 2026. It contains data from two East Asian cohorts, a Chinese cohort of 112,694 individuals and a Japanese cohort of 12,489 individuals, used to study the association between remnant cholesterol and incident diabetes. The analysis employed Cox regression models, restricted cubic splines, extreme gradient boosting, and mediation analysis.
2002 to 2014 data on enrolments in the UK Essential Skills programme, detailing subject areas, learner characteristics, and performance outcomes. The dataset is published by OpenDataNI under the OGL-UK-3.0 license and was last updated in July 2026.
OpenDataNI provides data on enrolments and performance for the UK Essential Skills programme from its start in 2002 through 2015/16. The dataset includes counts of enrolments by subject and details on the characteristics of enrollees. It was last updated on the platform in July 2026.
Brownfield sites within the Redcar and Cleveland planning area, regarded as suitable for development. The register is maintained by Redcar and Cleveland Borough Council and is intended to be updated annually, with a goal that at least 90% of sites would have planning permission by 2020.
Persons starting Community Punishment Orders, requiring 40 to 240 hours of unpaid work, or Rehabilitation Orders, combining 1-3 years of probation with 40-100 hours of community service, for individuals aged 16 or older. The dataset is a Data4NR reference from the UK Home Office, last updated in July 2026.
England and Wales administrative data on persons cautioned for summary (non-motoring) offences as a percentage of persons found guilty or cautioned. The dataset is broken down by police force area, sex, and age group. It was published by the Ministry of Justice and covers the period from 2000/01 to 2007/08.
Persons cautioned for indictable (excluding motoring) offences as a percentage of persons found guilty or cautioned. Ministry of Justice administrative data covers police force areas in England and Wales across eight fiscal years from 2000/01 to 2007/08.
Administrative data from the Ministry of Justice covering penalty notices for disorder issued to offenders of all ages in England and Wales. The data is aggregated by Police Force Area and covers the period from 2004 to 2007. It originates from the Ministry of Justice as the source and publisher.
OpenDataNI provides data on enrolments and performance for the Essential Skills programme in Northern Ireland. The dataset covers the period from the strategy's start in 2002 through the 2014/15 academic year. It includes enrolment numbers by subject and characteristics of those enrolling.
Butterfly records for Northamptonshire digitized from historical slips. The data covers the period from 1976 to 1985 and helps fill a gap in species records for the region. Records were digitized by the local branch of Butterfly Conservation under the NBN Data Capture Initiative 2014-15.
Percentage data of persons starting pre- or post-release supervision by the Probation Service, broken down by ethnic group and Police Force Area. The dataset was produced by the UK Ministry of Justice using administrative data. It covers England and Wales for the 2006/07 period.
Queensland resource authorities granted for activities like mining and energy. Reports are published by the Department of Natural Resources and Mines, Manufacturing and Regional and Rural Development when data is available for a period. The dataset was last updated on July 6, 2026.
A benchmark collection of 2,191 hypergraph instances originating from constraint satisfaction and query problems. The hypergraphs were generated and published by W. Fischl, G. Gottlob, D. M. Longo, and R. Pichler in 2017, with properties including various notions of width. Details on the original sources are provided in a 2018 conference paper by Johannes K. Fichte and colleagues.
Raw serial rotation electron diffraction datasets from two mixture zeolite products (Product A and B). The data includes calibration files, raw rotation diffraction data in SMV format, and on-the-fly unit cell identification results from DIALS. The dataset was authored by Yi Luo of Stockholm University and includes Python code for processing.
50 picoseconds of simulation data per trajectory, with frames saved every 0.0005 picoseconds. This dataset contains trajectories from ab-initio molecular dynamics simulations of TiO2 surfaces in water, as described in a 2017 Journal of Chemical Physics paper. The data was produced by Lorenzo Agosta of Stockholm University.
A dataset of labelled posts from Stack Overflow and discussions from GitHub Discussions related to GitHub Copilot. The data was collected by Beiqi Zhang of Wuhan University for an empirical study on the practices and challenges of using the AI coding assistant. It includes post IDs, URLs, and extracted content from both platforms.
Raw LC-HRMS data from two influent and effluent samples and a procedural blank, all in triplicate, used for benchmarking the patRoon software. The data was generated using an LC-Orbitrap Fusion instrument with positive ESI ionization. The dataset was shared by Dutch drinking water companies Dunea and PWN and authored by Rick Helmus of the University of Amsterdam.
Method chains collected from Java code for a study published at MSR 2020. The archive includes lists of chains, manual inspection results, and sampled data for specific research questions. Tomoki Nakamaru from The University of Tokyo authored this research dataset.
A dataset for numerical simulation using the snow algae model (Onuma et al., 2018; 2020). It contains algal cell concentration observed on surface snow worldwide, along with model input, output, and visualization scripts. The data and code were created by Yukihiko Onuma of The University of Tokyo.