Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,057 datasets
NACE Rev. 2.1 is the Statistical Classification of Economic Activities in the European Community. It is the European version of the ISIC and is harmonized with related classifications like CPC, HS, and SITC. The dataset is provided by Cooperation OGD Österreich and Wikimedia Österreich under a CC-BY-4.0 license.
NACE Rev. 2 is the European Union's standard classification system for economic activities. It is harmonized with the United Nations' International Standard Industrial Classification (ISIC) and replaced the previous NACE Rev. 1.1 version on January 1, 2008. The dataset is provided by Cooperation OGD Österreich and Wikimedia Österreich under a CC-BY-4.0 license.
CPA 2.1 is the European version of the CPC product classification, aligning products directly with their economic origin. It replaced the 2008 CPA version on January 1, 2015, and maintains a hierarchical structure analogous to NACE. The dataset is published under a CC-BY-4.0 license by Cooperation OGD Österreich and Wikimedia Österreich.
A 30-year timeseries of ocean observations from the eastern and southern coast of Australia underpins this statistical analysis of extreme storm events. The dataset includes multivariate summary statistics for storm events, such as maximum significant wave height, duration, and peak storm surge. This work is a key component of the Bushfire and Natural Hazards CRC Project Resilience to clustered disaster events on the coast.
Five distinct seabed sediment classes were identified in Keppel Bay, a macrotidal environment in Central Queensland, Australia. The Australian Ocean Data Network published this dataset, which combines sediment sampling with acoustic seabed mapping tools. Statistical techniques classified sediments based on grainsize, chemical composition, and modelled seabed shear stress from waves and tidal currents.
Bradley Jones authored a methodological paper proposing an alternative to Lenth's method for analyzing screening experiments. The work includes simulation studies comparing the new approach to current alternatives, with supporting files available in PDF, JPG, CSV, and TEX formats totaling 315.9 KB. The dataset was last updated on June 3, 2026.
3.0 KB of text codifies an industrial-grade, zero-allocation bare-metal processing framework. The registry, authored by Jamie Davis and licensed under CC BY 4.0, integrates a unified architecture for deterministic execution and high-resolution neuroanatomical mapping. It was last updated on June 1, 2026.
Q-mode factor analysis classified estuarine sediment samples from Broad Sound, Queensland, into two geologically distinct groups representing intertidal and supratidal deposition. The dataset likely contains measurements for variables including pH, Eh, P2O5, Cu, Pb, and Zn, analyzed using statistical techniques like discriminant and regression analysis. It was published by the Australian Ocean Data Network.
Gad Abraham developed a genomic risk score (GRS) model for celiac disease using statistical learning on genome-wide SNP profiles from six European cohorts. The model achieved cross-validation AUCs of 0.87–0.89 and independent replication AUCs of 0.86–0.9, explaining 30–35% of disease variance. This proof-of-concept demonstrates the potential for genomic-based risk stratification to improve clinical diagnostic pathways.
Radiocarbon data from archaeological cereal samples, compiled for a 2021 paper published in Radiocarbon. The dataset supports analysis of the emergence and development of arable farming in southeastern Norway. It was created by Steinar Solheim of the University of Oslo.
A methodological paper by Leo L. Duan from the University of Florida proposes a new algorithm for Bayesian posterior estimation. The method uses optimization to solve for a random transport plan between a posterior distribution and a simple uniform distribution. It is described as producing independent random samples with high approximation accuracy and is compared favorably to Markov chain Monte Carlo, variational Bayes, and normalizing flows.
BAMLSS is a unified modeling architecture for distributional generalized additive regression models (GAMs) established by Nikolaus Umlauf of Universität Innsbruck. It embeds many different Bayesian approaches from literature and software for capturing location, scale, and shape aspects of response distributions. The framework's usefulness is demonstrated with two complex application case studies: a large daily precipitation climatology and a Cox model for continuous time with space-time interactions.
Rafael Schmitt from UC Berkeley presents a stochastic model ensemble for sand connectivity in the Se Kong, Se San, and Sre Pok tributaries of the Mekong River. The model uses a Monte Carlo approach with the CASCADE framework to quantify uncertainty in sediment sources and upscale point observations to the entire network. The inverse stochastic approximation partitions sand deliveries, estimating fluxes of 1.9 Mt/yr from Se Kong, 5.3 Mt/yr from Se San, and 11 Mt/yr from Sre Pok.
A smartphone-based system for digital image colorimetric analysis using partial least squares regression. The method was validated by determining four adulterants in raw milk samples, achieving coefficients of determination higher than 0.99. The research was authored by Adilson Ben da Costa and published on Papers with Code.
RWTH Aachen University provides test instances for robust combinatorial optimization with budget uncertainty. The set contains nominal problems from MIPLIB 2017 converted into robust problems, instances of the robust knapsack problem, and instances for robust weighted matching and independent set problems. These instances were created by Christina Büsing, Timo Gersing, and Arie Koster for benchmarking algorithms described in published papers.
Daily climate data from 2015 to 2025 for Zhengzhou, China, including air temperature, relative humidity, wind speed, solar radiation, and precipitation. The dataset was created by Xiangrong Liu and last updated in May 2026 to support a framework integrating big climate data with differential equation models for environmental prediction. Forecasts derived from the model extend to 2030.
An adapted biventricular statistical shape model from Bai et al. (2015) provides a basis for computer simulations of cardiac electrophysiology and mechanics. The dataset includes two detailed meshes of the mean heart shape: a surface triangle mesh with 100k nodes and 200k elements, and a volumetric tetrahedral mesh with 479k nodes and 2.555 million elements. It also contains 100 principal components and variances for shape variation, plus 100 quasi-random shape instances generated by sampling weights within ±3 standard deviations.
Survey results from 2007 show the health of phytobenthos algae in English rivers, expressed against Water Framework Directive status classes. The dataset includes statistical information on the confidence that results have been correctly assigned to each status class. It was created by the Environment Agency and the assessment method has since been improved.
A study developed a method for determining calcium, magnesium, manganese, iron, and zinc concentrations in palm oil samples using flame atomic absorption spectrometry. The method, optimized via constrained mixture design, was applied to samples collected in Bahia State, Brazil, with reported detection limits between 0.012 and 0.057 mg L-1. Concentrations found in the samples ranged from 3.93-13.9 mg L-1 for Ca, 0.37-2.26 for Mg, >LOQ-0.32 for Mn, 1.77-8.57 for Fe, and 0.38-2.54 for Zn.
Philip S. Salmon from the University of Bath provides the data sets used to prepare figures for a Journal of Statistical Mechanics article. The data includes structure factors and bond angle distributions for materials like amorphous silicon, germanium, SiO2, GeO2, ZnCl2, and GeSe2. It supports analysis of ordering on different length scales in liquid and amorphous materials.