Loading...
Loading...
General ML benchmarks, tabular data, AutoML, recommendation systems, anomaly detection, evaluation suites
194,257 datasets
Mid-scale vector polygon dataset showing Ward boundaries from 1993. The data is published by OpenDataNI for OpenData under the OGL-UK-3.0 license. It was last updated on 2026-07-08 13:31:00.036720.
Northern Ireland's mid-scale vector polygon dataset shows administrative Ward boundaries from the year 1993. It is published by OpenDataNI for OpenData and is available in multiple geospatial formats. The dataset's license is OGL-UK-3.0, and its record was last updated on 2026-07-08.
OSNI 50k LGDs 1993, published for OpenData by OpenDataNI. The dataset was last updated on 2026-07-08 13:30:30.282816. By download or use of this dataset you agree to abide by the LPS Open Government Data Licence.
A paper from The University of Tokyo proposes a model and algorithm for discovering recurring partial periodic patterns in time series. The work addresses the challenge of finding patterns that exhibit periodic behavior only during specific intervals, which can provide information about seasonal or temporal event associations. The author is R. Uday Kiran.
Flow cytometry and 16S rRNA sequencing data support the development of machine-learned classification for microbiota cell type diversity. The dataset includes raw and cleaned measurements from synthetic microbial communities, E. coli cultures, and chemical enrichment experiments. It is supplemented by pre-trained artificial neural network functions and detailed methodological documentation.
Neural networks were trained on this dataset of curvature-optimal geometry parameters for smoothing robot motion paths. Benjamin Kaiser at the University of Stuttgart generated the data by offline solving an optimization problem for corner smoothing. The collection includes example trajectories for a Kuka KR500 R3330 robot, both smoothed and non-smoothed, with velocity planned under jerk limits.
A visual analytics system and demonstration data for interpreting hidden states in Long Short-Term Memory (LSTM) models. The project includes source code for preprocessing and visualization, with precomputed data for immediate use. It was created by Tanja Munz at the University of Stuttgart.
Supplementary material for a master's thesis from the University of Stuttgart. The data likely contains monitoring logs and event logs from a microservices experiment setup, including injected anomalies. The dataset includes raw monitoring data, anomaly logs, and results from different detection thresholds.
A limited archaeocyathan-radiocyathan fauna preserved in marine dolostones of central Australia includes species like Aldanocyathus greeni Kruse sp. nov. and Radiocyathus minor. The data originates from the Todd River Dolomite and Mount Baldwin Formation, part of the Early Cambrian platform cover. This dataset is hosted by the Australian Ocean Data Network and was last updated on 2026-06-23.
A qualitative research artifact from the University of Vienna contains source codings, an audit trail, and a generated Architectural Design Decision model for machine learning workflows. The dataset supports a Straussian Grounded Theory study of practitioner views documented in gray literature. It includes Python applications for generating results and a detailed description of the research method.
Two studies recorded eye movements while participants judged new job candidates after memorizing information about exemplars. The research, led by Bettina von Helversen at the University of Basel, investigated the activation of different information in memory depending on decision strategy. Results show longer fixations on locations of similar exemplars during similarity-based, but not rule-based, decisions.
175,706 automated software builds from two open-source CI/CD platforms spanning 10 years, collected by Lalit Narayan Mishra. The data supports a study on temporal data leakage in machine learning models for predicting build success or failure. Models using only pre-build features achieved 82.73% accuracy on TravisTorrent (2013-2017) and 83.30% on GHALogs (2023).
175,706 automated software build records from two open-source CI/CD platforms, TravisTorrent and GHALogs, spanning 10 years from 2013 to 2023. The dataset supports research on predicting build success while avoiding temporal data leakage, a methodological flaw identified by a three-type taxonomy. It was created by Lalit Narayan Mishra and is licensed under CC-BY-4.0.
175,706 builds from two open-source CI/CD platforms, TravisTorrent (100,000 builds, 2013–2017) and GHALogs (75,706 workflows, 2023), form the basis for evaluating machine learning models that predict build success. The dataset, created by Lalit Narayan Mishra and last updated in May 2026, provides metrics for a study on preventing temporal data leakage in build prediction. It demonstrates that removing leaky features reduces reported accuracy by up to 15.07 percentage points, revealing a more realistic performance baseline.
Aarhus University presents a large-scale electrical resistivity model database for deep learning applications in geophysics. It contains a wide variety of geologically plausible and geophysically resolvable subsurface structures for ground-based and airborne electromagnetic systems. The database aims to standardize data for benchmarking and evolving deep learning algorithms in this domain.
55,982 initially sampled Pt-Ni alloy structures with EAM energies and forces are provided in PtNi_alloy_eam.db. A final set of 6,828 resampled structures with DFT energies and forces calculated by VASP is available in PtNi_alloy_dft.db. Shuang Han from the Technical University of Denmark created this dataset for training neural network potentials and other machine learning models.
A method for efficient threshold queries on derived fields from large numerical simulation datasets stored in a relational database cluster. The datasets produced by these simulations are in the TB and PB ranges, making local analysis impractical. The approach achieves scalability through data-parallel execution and an application-aware cache that can improve query performance by over an order of magnitude.
1.5 million Monte Carlo simulated events evaluate momentum reconstruction for top quarks and their decay products. Jan Tuzlić Offermann from the University of Chicago produced this data, which includes versions with and without fast detector simulation using Delphes and the ATLAS detector card. Events simulate fully hadronic top quark decays at 13 TeV center-of-mass energy, with leading top quark pT between 550 and 650 GeV.
Experimental data from a study investigating how people map observed actions onto performed actions. The study involved participants performing individual or joint movements in synchrony with observed movements on a screen, with tempo increasing from 1.75 Hz to 3 Hz. The dataset likely contains results from multiple experiments showing differences in spatial accuracy between observing individual versus joint action.
Martha Larson of Radboud University Nijmegen argues for the principle of minimal necessary data in recommender systems. The paper illustrates the trade-off between training data volume and algorithm accuracy through a set of classic recommender system experiments. It concludes that adopting training data requirements analysis is a step towards more responsible recommendation.