Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,069 datasets
Teddy Lazebnik published a list of model variables for simulating ischemic dermal wound healing in May 2026. The dataset likely contains parameters for two mathematical models—a PDE and an Agent-based Simulation—used to assess oxygen therapy effectiveness. The models focus on keratinocytes, which constitute 90% of epidermal cells.
Michael Färber from Universität Innsbruck compiled proofs in the Dedukti format from multiple theorem provers. The dataset includes contributions from interactive theorem provers Matita, HOL Light, and Isabelle/HOL, as well as automated theorem provers iProver Modulo and Zenon Modulo. This collection serves as a benchmark for evaluating the proof checkers Kontroli and Dedukti.
Data replicates results from a paper submitted to IEEE Transactions on Evolutionary Computation. It compares Black-Box Optimization tools with classical heuristics on the BBOB benchmark suite from the COCO environment. The data was authored by Elena Raponi of Leiden University.
Jan-Peter Krämer from RWTH Aachen University provides the data required to reproduce statistical tests from the paper "An empirical study of programming paradigms for animation". The dataset supports the empirical findings reported in the study. Individual solutions produced during the study are available upon request.
Christopher Holder of Johns Hopkins University provides the dataset and scripts for the manuscript "Can machine learning extract the mechanisms controlling phytoplankton growth from large-scale observations? – A proof of concept study." The data is associated with research linking intrinsic and apparent relationships between phytoplankton and environmental forcings using machine learning. The dataset is released under an Open Access (green) license.
James D. Crall's data supports research on wind's role in pollinator dynamics in fragmented tropical forests. It includes orchid bee abundance, species identification, and site-specific forest cover and wind measurements. The dataset is accompanied by R scripts to reproduce the statistical analyses from the associated paper.
Milica Todorović from Aalto University created a configurational 5D DFT dataset of C60/TiO2 material. The dataset accompanies the manuscript 'Bayesian Inference of Atomistic Structure in Functional Materials'. The row count and specific file formats are unknown.
Sen Wang published a dataset on figshare in May 2026 detailing the isolation and characterization of a high-DHA mutant strain. The data likely contains results from a high-throughput Raman-activated cell sorting screen of 50,000 mutant cells and subsequent omics and fermentation analysis. It describes strain ABBS26, which achieved a final biomass of 147.3 g/L and DHA production of 47.4 g/L.
Kevin Wils's thesis repository contains codes and datasheets for a symbolic finite-element truss QUBO method. The data likely supports optimization of truss structures using quantum-inspired computing techniques. This work originates from Delft University of Technology.
2019 data and results from the article 'Reflectance spectra of seven lunar swirls examined by statistical methods: A space weathering study' by Chrbolková et al. The archive corresponds to source code, raw data, and results published in the Icarus journal. The dataset was authored by Kateřina Chrbolková from the University of Helsinki.
Research data was used for the article "Aššur and His Friends: A Statistical Analysis of Neo-Assyrian Texts" published in the Journal of Cuneiform Studies in 2019. The dataset was generated by Tero Alstola of the University of Helsinki. It likely contains statistical measures derived from ancient text corpora.
Giovanni Puccetti's pedagogical article examines common misconceptions about correlation in financial risk modeling. The work uses simplified examples and reproducible R code to demonstrate how misjudging statistical dependence can amplify systemic vulnerabilities, as seen in the misuse of the Gaussian copula during the subprime crisis. It aims to foster critical awareness of quantitative tools among risk managers, regulators, and students.
VIIRS/NOAA20 Deep Blue Level 3 monthly aerosol data provides a 1x1 degree gridded product derived from daily satellite observations. The dataset contains 45 Science Data Set layers in netCDF format, with monthly aggregation requiring at least 3 valid days of data per grid element. Its record starts from January 5th, 2018.
445.0 MB of data from a study combining multisource soil data, input-output analysis, and geospatial modelling to examine inequality in soil arsenic pollution burdens across China's provinces from 2000 to 2020. The work, authored by Ziyang Li and shared on figshare, shows the inequality index in recipient provinces rose 4.6-fold to 0.46, while sending provinces' index declined. Scenario simulations indicate land-use optimization could reduce overall inequality by about 25%.
5.5 KB of statistical model summaries for seven upper limb muscles, including m. brachioradialis and m. biceps brachii. The dataset, authored by Frida Torell and last updated in June 2026, contains results from multiple linear regression (MLR) models. It includes metrics like R² and Q², with a note on statistical significance thresholds.
Glenn T. Schumacher published a dataset on figshare on 2026-05-15 summarizing Bayesian hierarchical linear regression metrics from a stable isotope analysis study. The 5.5 KB XLS file contains results from an analysis of ontogenetic niche shifts in Arctic Charr and Brook Trout populations in four lakes in Maine, USA. The study used eye lens stable isotope analysis to investigate trophic position changes through life.
A de-identified clinical dataset collected for a study on Non-Invasive Prenatal Testing (NIPT). It includes maternal clinical information, gestational age, cell-free fetal DNA indicators, prenatal screening records, and confirmed fetal karyotype results for trisomies 21, 18, and 13. The data was used to construct multi-model algorithms for optimizing NIPT timing and predicting fetal chromosomal abnormality risk, matching a manuscript submitted to BMC Pregnancy and Childbirth.
L. M. Holanda presents a statistical mechanics approach to model magnetic material behavior. The work calculates magnetic susceptibilities and analyzes paramagnetic, ferromagnetic, and antiferromagnetic phases as a function of temperature. It is intended as a didactic resource for undergraduate students.
Performance measures for dataset ranking models, provided as complementary material for a research paper. The data was contributed by author Angelo Batista Neves Júnior. The repository includes results for 15 models, including Bayesian, cosine similarity, J48, JRip, and SN variants.
Archived website analytics for the Lincolnshire Open Data portal, split into three resources covering overall usage, summaries, and individual webpage statistics. The data was collected by the Government Digital Service and is licensed under the Open Government Licence. It specifically excludes API calls and includes only UK-based user traffic.