Loading...
Loading...
General ML benchmarks, tabular data, AutoML, recommendation systems, anomaly detection, evaluation suites
194,257 datasets
UNOSAT analysis FR20210811DZA maps vegetation burned by wildfires on August 12, 2021, in the Blida, Bouira, and Medea provinces of Algeria. The preliminary assessment, based on Sentinel-2 satellite imagery, estimates approximately 4,300 hectares of forest and vegetation cover burned across 240,000 hectares analyzed. The most affected communes were Boukram (1,350 ha), Deux Bassins (950 ha), and Souhane (650 ha).
Five types of multimodal annotated data cover the area around Chactún, a major Maya urban centre in the Yucatán peninsula. The dataset includes ALS visualisations, a canopy height model, Sentinel-1 SAR, Sentinel-2 optical data, and manual annotations for three structure types. It was published by Kokalj et al. in Scientific Data in 2023 and used for a computer vision competition.
Simulation output data from a particle-in-cell model investigating neutrino fast flavor instability. The dataset includes input parameters and timestamped output from a fiducial simulation, along with reduction and plotting scripts. Sherwood Richers from UC Berkeley produced this data using version 1.1 of the Emu simulation code.
Data from 2014 supports a study quantifying phytoplankton community structure using scanning flow cytometry and unsupervised clustering. The dataset contains raw flow cytometry measurements linked to environmental metadata like date, time, and sampling depth. It is intended for developing and validating machine learning methods to analyze aquatic microbial communities.
3,021 annotated document page images form the training, evaluation, and test sets for the ICDAR 2019 Competition on Baseline Detection (cBAD). The dataset was created by Markus Diem of TU Wien and consists of real-world images collected from seven European archives, with all baselines manually annotated. The training and evaluation sets contain PAGE XML files with annotated text regions and baselines.
A dataset of 3,021 annotated document page images collected from seven European archives for the ICDAR 2019 Competition on Baseline Detection (cBAD). All baselines were manually annotated, and the training and evaluation sets include PAGE XMLs with annotated text regions and baselines. The dataset was created by Markus Diem of TU Wien.
A study by Ithalo Coelho de Sousa used genomic data from 245 Arabica coffee plants genotyped for 137 markers to predict resistance to orange rust. Machine learning algorithms, including Decision Trees and their refinements, Artificial Neural Networks, and Bayesian Generalized Linear Regression, were compared for prediction accuracy and marker importance identification. The refinements identified an average of 9.3 important markers located in quantitative trait loci regions associated with disease resistance.
An R implementation of a spectral clustering algorithm proposed by David P. Hofmeyr. The algorithm addresses the practical challenges of selecting the number of clusters and tuning the scaling parameter by leveraging the asymptotic value of the normalised cut. It is available as an open-source package on GitHub.
All raw data, intermediate steps, and final results for three experiments used in the study 'Strong El Niño events lead to robust multi-year ENSO predictability'. The data has been gridded to a common 2x2 degree grid. Nathan Lenssen from the University of Colorado Boulder authored the dataset, which is archived alongside its codebase.
CoCO2 project data reconciles observation- and inventory-based methane emissions for eight major emitting countries. The spreadsheets contain CH4 data behind manuscript figures, presented as time series, mean values, and uncertainty ranges. This synthesis is based on data from the VERIFY project for the EU27 and deliverables from the CoCO2 project funded by the European Commission.
Mid-infrared spectroscopy chemical images of a breast cancer tissue microarray, specifically sample BR20832 from Biomax. The data relates to an open-access paper published in Analyst by researchers from the University of Manchester exploring machine learning for infrared pathology. Processed versions of the data are available in a separate Zenodo archive.
A method for fatty acid determination in fermented milk reduces sample and reagent use while shortening experiment time. Validation on yogurt samples with fat content from 1.22% to 2.83% showed intraday and interday relative standard deviations of 1.86% and 4.23%, respectively. The dataset, associated with a paper by Angélica F. B. Piccioli, likely contains results from applying this method.
Data collected from the 1970s to present provides physical and chemical properties, sedimentary processes, and glacial and marine history of the terrestrial environment in the Vestfold Hills, East Antarctica. This compilation incorporates samples and observations from published and unpublished sources, presented as point locations. The Australian Ocean Data Network hosts the data.
NHS Pennine Care's expenditure data details individual transactions over £25,000 from January 2019. The dataset is published under an open CC-BY-4.0 license by the UK Government Digital Service, indicating a commitment to public financial transparency. Its cross-platform presence on UK and EU open data portals suggests it is a recognized public finance record.
February 2019 expenditure data from Pennine Care NHS Foundation Trust details individual transactions exceeding £25,000. This dataset provides a granular view of public spending within a specific UK healthcare provider for a single month. Its cross-platform presence on UK and EU open data portals signals its use for public accountability and financial analysis.
Measured potential profile data accompanies the manuscript 'Measured potential profile in a quantum anomalous Hall system suggests bulk-dominated current flow.' The dataset and analysis code were created by Ilan T. Rosen of Stanford University and are available via the arXiv preprint server and a GitHub repository.
Raw electron diffraction data for the as-synthesized zeolite SSZ-27, including data for an SSZ-26 impurity phase. The data were collected using the software instamatic and processed using XDS/edtools by Stef Smeets of Delft University of Technology. It includes raw image files, processing logs, and clustering results from multiple crystals.
Raw data underlying published and unpublished figures from the retracted manuscript "Quantized Majorana conductance" by Zhang et al. from Delft University of Technology. The repository includes analysis methods and a side-by-side comparison between original and corrected figures. A public version of an expert report commissioned by TU Delft is also referenced.
Global coverage of river flood susceptibility across 8 level 1 drainage basins. The dataset contains two raster maps per basin: one for upstream accumulating area and one for Strahler stream order of flooded rivers. It was produced by Mark Bernhofen of the University of Leeds using the methodology from Bernhofen et al. 2021.
A video flythrough presents seabed bathymetry compilations for the Australian Antarctic margin. The data is derived from multibeam, singlebeam, and satellite sources, including ETOPO2. The Australian Ocean Data Network published this resource, which was last updated on 2026-06-23.