Loading...
Loading...
General ML benchmarks, tabular data, AutoML, recommendation systems, anomaly detection, evaluation suites
192,812 datasets
Southern Ocean waters south of Tasmania were sampled during the RV Investigator Eddy voyage IN2016_V02. The dataset contains water samples from mesoscale cyclonic and anticyclonic eddies, analyzed for phytoplankton biomass and nutrient concentrations. It is hosted by the Australian Ocean Data Network and was last updated in July 2026.
North East England's 2011 Census output areas are defined in this geospatial PDF map. The dataset provides the official administrative boundaries used for statistical reporting and demographic analysis in the region. It is sourced from the UK Government Digital Service and is available under an Open Government License.
Lagrangian trajectory data simulated for 2000 model years using the TRACMASS v7.0.0 code and EC-Earth/NEMO model output. The dataset contains initial, end, and full trajectory positions for particles released in the Drake Passage, run both forward and backward in time from 1986 to 2005. Sara Berglund from Stockholm University created this dataset, which loops the 20-year monthly mean ocean state to extend simulations.
Benchmark data sets and pre-generated input features for the AF2Complex deep learning method, focusing on protein complex prediction. The data includes benchmark sets CP17, Dimer1193, and Oligomer562, as well as features for 4,429 proteins from the E. coli proteome. This dataset was created by Mu Gao of the Georgia Institute of Technology to support the work 'Predicting direct physical interactions in multimeric proteins with deep learning'.
Silvan Sievers from the University of Basel provides code, scripts, and data for reproducing experiments from the SoCS 2022 paper on additive pattern databases. The bundle includes implementations based on Fast Downward, experiment scripts for Lab 7.0, and benchmark suites from IPC and Autoscale. Raw and processed experimental data from planner runs are also included.
Monthly Ocean Heat Content Anomalies (OHCA) in the top 2000 dbar of the North Atlantic ocean north of 20°N are calculated for the period 2005-2022. The data is mapped using a locally stationary Gaussian process from Argo float observations, with separate vertical sections from 15 to 1850 dbar. Regions shallower than 300 m or not well-sampled by Argo are excluded.
The paper describes a machine learning toolkit tested on six distinct datasets. The datasets were used to evaluate the MLme tool, which integrates Data Exploration, AutoML, CustomML, and Visualization functionalities. The supporting data originates from the University of Bern and accompanies an open-source toolkit for classification problems.
A database of Hidden Markov Models (HMMs) for protein domains, mapped from the TED database to CATH structural classifications and filtered for a maximum pairwise sequence identity of 30%. Claudia Alvarez-Carreño authored this dataset, which was last updated on June 4, 2026. It contains 934,186 HMMs spanning 4,688 CATH superfamilies.
Etan Green of Stanford University created a dataset analyzing over a million decisions made by Major League Baseball umpires. The data tests propositions about incentives solving principal-agent problems and reducing decision biases. It reveals a systematic aversion among umpires to choices that would more strongly change a game's expected outcome, termed 'impact aversion'.
Two gzipped tarballs contain BED files of Virtual ChIP-seq posterior probabilities for transcription factor binding. The data corresponds to predictions on specific chromosomes (chr5, chr10, chr15, chr20 for Cistrome; chr1, chr8, chr21 for ENCODE-DREAM) in validation cell types. The dataset was created by Mehran Karimzadeh of the University of Toronto.
Bradford Council's dataset on counter fraud work is published in compliance with the UK's Local Government Transparency Code 2015. The dataset is published by the City of Bradford Metropolitan District Council and was last updated on 2026-07-08. Its specific content and scale are not detailed in the available metadata.
16 December 2025 notices for fiber optic infrastructure work in Cheltenham, Victoria, Australia. The dataset is published by the SIP Register - Fiber Asset Management Pty Ltd on data.gov.au under a CC-BY-4.0 license. It includes file formats such as ZIP MAPINFO and .PDF, suggesting geospatial and document data.
Victorian Property Sales Report - Median House by Suburb Quarterly is published by the Department of Transport and Planning on data_gov_au. It lists the percentage shift in median property prices between quarters and over a 12-month period, including overall metropolitan and country Victoria medians for each property type. The dataset was last updated on 2026-07-01.
Observations of sea level at Port Arthur, Tasmania during the period 1840 to 1842. The data includes date, time, and sea level measurements from a tide gauge constructed by T.J. Lempriere, with observations made from mid-1837 to at least the end of 1842. The dataset is provided by the Australian Ocean Data Network.
A predictive model dataset for annual hard coral cover across all known coral reefs in the Central Great Barrier Reef and American Samoa. The data, aggregated into 5x5km hexagonal units from the public ReefCloud platform, incorporates exposure to heat stress and tropical cyclones. This repository supports the paper 'Predicting coral cover trends from local to broad spatial scales' by Vercelloni et al. (2026).
Office for National Statistics provides an exact fit lookup file between Middle layer Super Output Areas (MSOAs) from December 2011 and December 2021, and Local Authority Districts from December 2022 in England and Wales. The dataset includes a 'change indicator' field with four categories to define changes between the 2011 and 2021 MSOA boundaries. This version 2 includes updates to the change indicator for splits that went to complexes in under 10 MSOAs.
The 2017 Northern Ireland Multiple Deprivation Measures (NIMDM 2017) were published on 23rd November 2017 by OpenDataNI. It provides information for 7 distinct types of deprivation, known as domains, along with an overall multiple deprivation measure, comprising 38 indicators in total. The measures rank areas within Northern Ireland from most to least deprived but do not quantify the extent of deprivation differences.
A one-off publication of Bristol City Council's existing waste collection contracts, required by the Local Government Transparency Code 2014. The data is published by Bristol City Council and is available in CSV and HTML formats. Quarterly updates are available from a separate Bristol City Council Contracts and Tenders dataset.
Three real-world energy datasets were used to evaluate a novel forecasting model called DG-LSTM-SA. Guoqiang Sun published these data statistics on figshare in June 2026. The datasets likely contain time-series records of power generation and load demand.
Hyperparameters for ten baseline models evaluated in a study proposing a DG-LSTM-SA network for power generation and load demand forecasting. The dataset, authored by Guoqiang Sun and uploaded to figshare, is a 5.5 KB Excel file last updated on June 3, 2026. The models were tested on three real-world energy datasets: NEPOOL, Yichang, and Solar-Energy.