Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,060 datasets
A study by Benjamin Smith of the University of Florida examines endogenous selection bias in cross-national statistical research. The research uses a causal model of post-colonial sovereignty on the Arabian Peninsula to illustrate survivorship bias, showing the effect of oil on autocratic survival becomes negligible when corrected for bias. The work motivates context-sensitive statistical modeling for causal inference with observational data.
Raw cross-hole time-lapse Electrical Resistivity Tomography (ERT) data and inversion results from a field study in Sardinia. The data package supports research into freshwater-saltwater interactions in porous media, with applications for underground freshwater storage. It was produced by Klaus Haaken of the University of Bonn and accompanies a published scientific manuscript.
Experimental and synthetic image data for structured illumination microscopy (SIM) reconstruction. The collection includes fluorescence images of calibration slides and live HeLa cell mitochondria, synthetic line pair patterns, and camera calibration maps. The data was associated with a research paper by Ayush Saurabh of Arizona State University.
A U.S. sample of 1,130 children aged 6 to 16 years who were assessed for learning difficulties provides data for this Bayesian structural equation modeling study. Author Philippe Golay from the University of Geneva analyzed the Wechsler Intelligence Scale for Children (WISC-IV) core subtest scores to compare factor models. The analysis found support for a bi-factor model over a higher-order model, suggesting a simple interpretation of the subtest scores.
A dataset containing performance evaluations and rankings for second pillar pension funds in the Baltic States. The analysis uses a Network Stochastic Dominance (NetSD) ratio to rank funds under various distributional assumptions, including empirical, α-stable, Student’s t, hyperbolic, and Normal Inverse Gaussian. The dataset was created by Audrius Kabašinskas and last updated on June 2, 2026.
Hugo Banderier from the University of Bern created this dataset for the manuscript "Reduced floating-point precision in regional climate simulations: An ensemble-based statistical verification". It contains spatially averaged test results for five experiments: Identical, Single Precision, and three Modified Diffusion experiments. The data includes results for ten years and 100 random selections, along with final verification decisions based on a 95th percentile threshold.
Benchmark programs used in the POPL'24 paper "Commutativity Simplifies Proofs of Parameterized Programs" by Azadeh Farzan, D. Klumpp, and A. Podelski. The dataset is associated with a paper published in the Proceedings of the ACM on Programming Languages (POPL) in 2024. The archive is provided by the authors from the University of Toronto.
Quantitative RT-PCR data are analyzed using generalized linear mixed models based on lognormal-Poisson error distribution, fitted using MCMC. The package implements a lognormal model for higher-abundance data and a classic multi-gene normalization model. Mikhail V. Matz of The University of Texas at Austin authored this package, which includes several plotting functions for result visualization.
Live tables from the UK government provide the latest and most popular data on housebuilding. The spreadsheets are primarily produced from statistical returns completed by Local Authorities, with some data from surveys or external sources. The Ministry of Housing, Communities and Local Government maintains this data, which was last updated on 2026-07-08.
Jamie Davis published a 7.5 KB dataset on figshare in 2026, consisting of a mathematical simulation matrix and C++ code for signal damping. It models predictive stabilization of corrupted telemetry feeds under chaotic noise conditions. The repository includes three high-interference simulation scenarios for deep space, subsea, and urban smart-grid environments.
617 community-dwelling participants aged 65+ were randomly assigned to upper or lower limb exercise groups in a 12-month trial. The dataset includes primary outcomes measured by the Disabilities of the Arm, Shoulder, and Hand (DASH) questionnaire and secondary outcomes like shoulder strength, range of motion, quality of life, and physical activity. Amanda Bates authored this dataset, which was last updated in June 2026.
A 2026 reference dataset by independent systems architect Jamie Davis, providing two structural infrastructure layers for the Davis Logic V2 bare-metal framework. It includes an invariant asymmetric slew-rate limiter engine and a cache-line contiguous memory alignment specification. The dataset is a 4.3 KB text file licensed under CC BY 4.0.
Davis Logic V2 is an open-access reference dataset providing an immutable, real-time hysteresis state debouncer core for bare-metal frameworks. It includes a branchless integer calculation engine and an automated testbench script to simulate unstable electrical connections. The dataset was authored by Jamie Davis and published on figshare in 2026.
A cross-sectional study of 93 children under one year assesses the prevalence of iron-deficiency anemia and vitamin A deficiency. The study found a 29.03% prevalence of anemia and a 19.10% prevalence of vitamin A deficiency, with low vitamin A values in 90.32% of children. It was authored by Mariane Alves Silva and uses statistical analysis from Stata 10.0 software.
A research dataset from paperswithcode evaluates an acoustic-based method for isolating human platelets from whole blood. The study compares platelet purity, activation status, and functionality between acoustophoresis and a reference soft-spin centrifugation protocol. Results indicate acoustophoresis eliminates over 80% of red blood cells and leukocytes, achieving a mean platelet purity of 92.8 ± 12.8% with minimal activation.
A study of 36,707 intoxication cases attended to by a poison control center in Brazil, of which 8,608 (22.5%) were drug-related. The data, collected from service notification records, was analyzed using linear regression to identify trends over a 30-year period. The research was authored by Thays Lopes Mathias and published via an open-access platform.
A quantitative cross-sectional study of 23 post-menopausal women, analyzing acoustic voice parameters to investigate vocal aging. The data was collected using Praat software and includes measurements of fundamental frequency, jitter, shimmer, harmonic noise, and spectrographic analysis. The study was conducted by researcher Renata D’Arc Scarpel with participants from the Open University of Third Age in Salvador, Brazil.
Farah Assadian's research paper describes a dataset likely containing spectrofluorimetric measurements for determining ofloxacin concentration in spiked human urine samples. The dataset likely includes calibration data spanning 5.0-120.0 ng mL-1, with results from multivariate calibration models like OSC-PLS and OSC-PARAFAC. The data was generated using dispersive liquid-liquid microextraction and Box-Behnken design optimization.
81.15% maximum aglycone yield was achieved under optimized conditions of 2.5 M HCl at 75 ºC for 60 minutes. The dataset likely contains experimental results from applying response surface methodology to optimize the hydrolysis of myricetin-3-O-rhamnoside and Inga edulis leaf extract. The study was conducted by Tatiana Tolosa and published via an Open Access platform.
72 snap bean accessions from the Universidade Estadual de Londrina germplasm bank were analyzed using amplified fragment length polymorphism (AFLP) markers. The study characterized 40 determinate and 32 indeterminate growth habit accessions, revealing 485 informative loci. Analysis by Felipe Aranha de Andrade showed indeterminate accessions exhibited greater variability and suggested associations with Andean and Mesoamerican gene pools.