Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,080 datasets
18.9 GB of processed Ethereum blockchain transaction records for ERC20 tokens, supporting a study on statistical patterns. The dataset, created by Kundan mukhia, provides a structured subset of transaction-level data categorized by interaction types. It was last updated in April 2026.
FinProof v1 is an open adversarial benchmark designed to test AI guardrail systems in banking, financial services, and insurance. It covers 7 attack categories across professional and retail conversational registers and was published by Zytra under a CC BY 4.0 license. The dataset was last updated on Hugging Face in May 2026.
Barak Hadad's dataset contains 938.9 MB of processed neural data files required to reproduce results from the study "Auditory network persistence of stimulus representation in awake and naturally sleeping mice." It includes pre-processed neural activity, decoding outputs, and statistical summaries, last updated on April 27, 2026. Raw electrophysiological recordings are not included but are available upon request.
623 stream sites in southwestern Yukon provide geochemical data for 36 elements in sediments, plus water measurements for uranium, fluoride, and pH. This dataset was compiled by the Government of Yukon and published as GSC Open File 2859/EGSD Open File 2001-11(D). The data was last updated on the platform in April 2026.
Beginning in 2006, this dataset tracks monthly counts of approved public assistance applications across New York State's Local Social Services Districts. It provides separate totals for Family Assistance and Safety Net Assistance case openings, mirroring annual statistics published in official state reports. The data is maintained by the New York State Office of Temporary and Disability Assistance.
Nigeria conflict data from 1997 to 2023 combines spatial econometrics with a greed-grievance framework. The analysis includes 14 variables across four types of conflict, using Bayesian spatial methodology and INLA-SPDE techniques. The dataset was authored by Juan José Villar-Roldán and is hosted on the Political Science Research Methods Dataverse.
Geoscience Australia Data provides a theoretical framework for representing isostatic processes using mathematical filters called admittance functions. The development of cross-spectral techniques relates gravity and topography, with examples for elastic and visco-elastic rheologies. The last update was recorded as 2026-04-20 01:34:17.675553.
A textbook titled 'The Complete Practical Arithmetician' containing arithmetic improvements for educational use. The dataset is published on the paperswithcode platform. The original author, publication date, and specific content details are unknown.
Statistical results from datasets covering multiple cell types, including bladder, kidney, frozen and fresh tumor, and mouse cortical cells. The dataset was authored by Xiran Chen and is available under a CC-BY-4.0 license. It was last updated on May 6, 2026.
Feng Li's dataset provides statistical evaluation metrics for models simulating soil moisture and salinity. The dataset is stored in an XLS file with a size of 5.5 KB and was last updated on May 6, 2026. It is licensed under CC-BY-4.0 and hosted on the figshare platform.
62,555 dementia cases and 312,772 matched controls from nationwide Finnish health registries were analyzed to examine the role of 29 hospital-treated diseases in the association between severe infections and dementia. The study, authored by Pyry N. Sipilä and published on figshare, identified that the increased dementia risk from two infectious diseases was not attributable to 27 other comorbid conditions. The dataset, last updated in March 2026, contains the statistical codes used for this analysis.
520,000 error traces document models' mistakes during math problem synthesis. The dataset includes updates from March 2026, where 50,000 new datapoints were added and 10,000 older ones were replaced with higher-quality synthetic questions verified by a 12-consensus tool. It was authored by nguyen599 and last updated on Hugging Face in May 2026.
Linear Discriminant Analysis results probabilistically assigning volcanic clast samples to sources between the Villa Draghi and Via Scagliara di M. Castellone quarries. The dataset, created by Simone Dilaria, is a 5.5 KB Excel file last updated in April 2026. Discriminant functions with a p-value < 0.05 are considered statistically significant at the 95% confidence level.
Simone Dilaria's dataset contains the results of a Linear Discriminant Analysis (LDA) performed on volcanic clast samples. The analysis probabilistically assigns samples to sources between the Villa Draghi and Via Scagliara di M. Castellone quarries, using a discriminant function considered statistically significant at a 95% confidence level (p-value < 0.05). The dataset was last updated on April 13, 2026.
Monthly updated statistics on hotel stays, guests, and occupancy rates in Utrecht's public area. The dataset, sourced from CBS Statline and managed by Utrecht Marketing, also includes reports on tourism, business visits, conferences, and monthly visitor figures for museums, theaters, and historical buildings behind a login. It contains descriptive and statistical information about residents, visitors, companies, and talents in the city.
Statistical data on licensed ferment-on-premises operators in Nova Scotia. The dataset compares the total volume of production of wine and beer, excluding kits sold for home production. It is provided by the Government of Nova Scotia and was last updated on April 17, 2026.
26.9 hours of motion data from 140 subjects across 10,386 trials. ForceBody pairs the SKEL parametric body model with measured ground reaction forces and inverse-dynamics joint torques. A subset of 8,652 trials includes per-frame, per-joint Monte Carlo uncertainty estimates for torque labels.
OR-Space is a full-lifecycle workspace benchmark for evaluating LLM agents on industrial optimization tasks. The benchmark, created by Chenyu-Zhou, structures each instance with separate files for business requirements, parameters, source code, and solver artifacts. It was last updated on May 18, 2026.
384 hours of observations with the Karl G. Jansky Very Large Array (VLA) at 3 GHz produced this catalog of 10,830 radio sources down to a 5-sigma threshold. The survey covers the 2 square degree Cosmic Evolution Survey (COSMOS) field with a median rms of 2.3 μJy/beam at 0.75 arcsecond resolution. The catalog was created by the HEASARC in June 2017 based on the reference paper's data release.
Shijuan Yang's dataset contains lifetime data for gas turbine components, including failure times and controllable factor settings. The 34.1 KB XLSX file is used to demonstrate a mixture Weibull regression and multi-objective optimization framework for lifetime-oriented quality design. It was last updated on 2026-04-30 and is shared under a CC-BY-4.0 license.