Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,068 datasets
Experimental data on the climbing distribution of 60 day-old male Drosophila flies expressing human SOD1 proteins in different cell types. The dataset likely contains counts of flies in three height bins within a glass cylinder, comparing expression in motoneurons, glial cells, and dual expression. The data was generated by Rafique Islam and is available via paperswithcode under an Open Access license.
A dataset associated with a 2017 Optics Express submission by Jean-François Rupprecht and colleagues. It contains sequences of localization events from labelled tubulin proteins, gold particles, and functionalized silane coverslips. The data was used to analyze trade-offs in stochastic super-resolution microscopy techniques.
A collection of images captured to align a spectrograph sensor in its best-focus position. The dataset includes all computer code to compute the best-focus position and determine confidence intervals for the solution, providing a statistical measure of precision. It was authored by Travis W. Sawyer and includes an example IPython notebook for guidance.
Rand R. Wilcox provides data, code, and notebooks to reproduce all analyses and figures from the 2018 paper 'A Guide to Robust Statistical Methods in Neuroscience'. The package includes R functions and PDFs of all figures, along with extra notebooks describing simulations on statistical power and type I error. The work describes and illustrates improved methods for comparing groups and studying associations.
An empirical model developed from a 2^3 factorial design enables control of hydroxyapatite nanoparticle shape and size. The model was used to synthesize nanoparticles with sizes between 8 and 600 nm, confirmed by TEM and SEM imaging. The research was conducted by Thaís M. Arantes and published as an Open Access paper.
Wei Zhang's dataset supports a study on variational inference for Bayesian dynamic generalized additive models applied to mortality analysis. The data includes Italian death counts from January 2015 to December 2020, with predictors and temporally dependent multivariate random effects. The associated files include PDF, RDATA, TXT, STAN, R, and CPP formats, totaling 2.1 MB.
A dataset likely contains experimental results from a medicinal chemistry optimization campaign targeting Rho-associated coiled-coil containing protein kinase 2 (ROCK2) inhibitors for pulmonary fibrosis. The data was uploaded by Zhi Cao to figshare on 2026-05-25. It includes measurements such as IC50 values, kinetic solubility, metabolic stability, and in vivo efficacy for a series of compounds.
Historical statistical data from Major League Baseball (MLB) compiled from Stathead. The dataset covers the period from 1904 to 2025. It was contributed by Ricardo de la Peña to the Harvard Dataverse and was last updated in July 2026.
Eight regional boundaries defined by the Aboriginal Affairs Coordinating Committee for reporting COAG Closing the Gap indicators in Western Australia. The geometry is based on ABS Statistical Areas SA2, simplified to reduce file size, and is compatible with Regional Development Commission boundaries. The dataset was last updated on 2026-06-29 by author Chris Dorrian.
Michael A. Irvine published a mathematical modeling study on figshare in 2026. The dataset likely contains parameters and results from a stochastic SEIR-like model calibrated to measles outbreaks in two regions of Northern British Columbia, Canada, during 2025. It was used to simulate scenarios involving public health interventions, mobility, and school terms.
236 patients with Graves' disease received initial single-dose radioactive iodine therapy between August 2022 and August 2024 at a tertiary center in Thailand. The dataset likely contains thyroid volume categories, administered activity per gram, and 6-month outcomes including remission and adverse events. Author Keattichai Keeratitanont published this retrospective study under a CC-BY-4.0 license.
A Bayesian network meta-analysis synthesizing 39 randomized controlled trials (N=2,714 participants; mean age 58.8 years) to compare exercise modalities. The analysis, authored by Gang Huang and last updated in April 2026, evaluates the efficacy of aerobic, resistance, high-intensity interval, and combined training on chronic low-grade systemic inflammation biomarkers in postmenopausal women with overweight or obesity.
2017-2021 statistical information for the Ruse region in Bulgaria. The data was compiled by the Data Department of the State e-Government Agency and is available via the eu_open_data platform. The specific variables and volume of data are not detailed in the available metadata.
Aggregated data from scans of Carers' Allowance data at 1992 ward level. The dataset is published by OpenDataNI under the OGL-UK-3.0 license. It was last updated on 2026-07-08 13:33:05.581840.
A subset of STRIPS planning problems from the sequential optimization tracks of the International Planning Competition (IPC) held between 1998 and 2018. The dataset is sourced from the AI research group at the University of Basel and is available via the paperswithcode platform under an Open Access license. It likely contains formal problem definitions used for evaluating automated planning algorithms.
A 2026 Bayesian network meta-analysis by Yi-Ming Li, synthesizing data from 54 randomized controlled trials. It models the nonlinear dose-response relationship between exercise and lower-limb outcomes in middle-aged and older adults with stroke. The study identifies an optimal total exercise dose of approximately 1,800–2,200 MET×min/week.
The dataset and R-code by Emma Anna Roosje Zuiderveen support analysis of bio-based products' potential to reduce environmental impacts. The R-code includes models for predicted mean relative risks of greenhouse gas emissions and other impacts, with arithmetic averages and 95% confidence intervals. It also contains fixed-effect moderator models for further statistical investigation.
Australian Department of Social Services payments data aggregated by 2021 Statistical Area 2 (SA2) for use in National Map. The data is privacy-protected with cell rounding to the nearest 5 from December 2022, and prior to that, values under five were randomly assigned. This dataset was last updated on May 15, 2026.
43 transpositional tetrad templates form the core of this supplemental data for the ChiRho transformational framework in music theory. The dataset includes catalog tables, operator definitions, orbit data, and musical examples supporting the manuscript by Bryan Roberts, last updated in May 2026. Files are provided in formats including PDF, TXT, and CSV.
Module 67 provides a real-time signal reconstruction mechanism for approximating missing telemetry or sensor data frames. The 2.5 KB TXT file contains the production C++ source implementation for a deterministic fixed-point arithmetic model. Authored by Jamie Davis and released under CC BY 4.0, it was last updated on May 30, 2026.