Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,064 datasets
A study by Gislayne Christianne Xavier Peixoto compared protocols for estrus induction in agoutis (Dasyprocta leporina). Ten female agoutis were treated with either cloprostenol alone or cloprostenol associated with a GnRH analogue, with reproductive cycles monitored via blood tests, ultrasound, and vaginal cytology. The study concluded that both protocols showed limited efficiency, only inducing estrus in females that were in the luteal phase at the start of treatment.
A cross-sectional study of 46 elderly adults aged 60 years and older from Anápolis, Brazil, evaluated visual functions and their relationship to functional vision and falls. The research, authored by Amanda Alves Lopes and published via paperswithcode, found statistically significant correlations between stereopsis and self-reported falls, and between visual acuity and functional vision.
A randomized controlled trial involving 65 adults with type 2 diabetes mellitus, with 12 months of follow-up, to evaluate the effect of an implementation intention strategy on promoting walking. The intervention group showed a statistically significant increase in leisure-time physical activity and a decrease in waist circumference compared to the control group. The dataset likely contains anthropometric measurements and survey instrument results from this trial.
Posterior distribution samples of model parameters from the paper "A stochastic intracellular model of anthrax infection with spore germination heterogeneity". The dataset includes Python computer codes to generate Figures 4, 5, 6, 7, 8, and 10 of the paper, which was published in Frontiers in Immunology. The author is Bevelynn Williams from the University of Leeds.
Area of statistical districts in Linz, Austria, measured in hectares. The dataset is published by Cooperation OGD Österreich and Wikimedia Österreich on the eu_open_data platform. It is available in CSV and ODS formats under a CC-BY-4.0 license.
Monthly statistics on clearances of cigarettes and tobacco products and duty receipts for the UK, produced by HM Revenue and Customs. The data is designated as National Statistics and is published in HTML format. It was last updated on 2026-06-18.
870 expert-validated editing triplets form the core of ChartSync, a benchmark for evaluating Visuo-Logical Cascading Editing (VLCE) in statistical chart images. The dataset, created by JiakangYu, includes 9 chart categories and 4 task types, with 235 geometry-coupled VLCE instances. It was last updated on July 11, 2026.
A statistical report covering boating incidents in New South Wales for the 10-year period ending 30 June 2018. The dataset is published by the NSW Government on the data.gov.au platform. The report is available in PDF format and was last updated on 2026-06-26.
A 2026 analysis by Rika Anderson identifies clusters of orthologous genes (COGs) whose abundances correlate with nutrient concentrations in the global ocean. The dataset is derived from 139 Tara Oceans metagenomic samples, analyzing 4,787 COGs against environmental metadata including phosphate, nitrate/nitrite, oxygen, and modeled iron. Statistical models were applied to control for confounding effects from variables like temperature, depth, and salinity.
A GPU-accelerated topology optimization pipeline written in Julia and CUDA for large-scale 3D structural design. The code implements the SIMP formulation with Heaviside projection and a matrix-free multigrid-preconditioned conjugate gradient solver, and includes an optional porous infill extension validated on a femur reconstruction case at approximately 77 million degrees of freedom. It was authored by Vadillo Morillas, Abraham, and last updated on June 21, 2026.
London local authorities' actions under homelessness provisions from the 1985 and 1996 Housing Acts, covering the financial years 2004/05 to 2017/18. The dataset is compiled from the DCLG P1E Homelessness returns and is maintained by the Greater London Authority as a measure of Economic Fairness.
Chilean women aged 50–81 years from the Magallanes Region were assessed for cognitive performance across three climacteric stages. The dataset, created by Jonathan Lühr-Henríquez and last updated in 2026, contains results from the Addenbrooke’s Cognitive Examination-Revised and Symbol Digit Modalities Test for 360 participants. It supports a Bayesian multivariate analysis of how the relationship between age and cognition varies by menopausal stage.
Chinese short-form dramas on YouTube are analyzed in a dataset of 895 videos uploaded between April and October 2025. The data, uploaded by Yunqiu Yan, includes channel subscriber base size, video duration, and upload timing to model predictors of cumulative view counts. Linear regression, multilayer perceptron, and random forest models were used, with the random forest achieving an R² of 0.666.
A dataset from figshare describes compounds identified as inhibitors of YAP-TEAD-dependent transcription. The data, authored by Timo Heinrich and last updated in May 2026, includes results from a TEAD-reporter-based cellular screen, thermal shift assays, and in vivo xenograft studies. Derivatives with varied selectivity profiles and optimized physicochemical properties were generated.
Marzia Bordone from the University of Zurich created this archive of pseudo events for the decays Bbar -> P mu X_nubar, where P is a D or pi meson. The data complements the publication EOS-2016-01. Events are stored as binary files in the HDF5 format.
Hua Wang authored a paper introducing the Edgeworth Accountant, an analytical method for composing differential privacy guarantees. The method leverages the f-differential privacy framework to track privacy loss under composition, providing non-asymptotic (ε,δ)-differential privacy bounds. The associated files, last updated on 2026-05-18, are available on figshare under a CC-BY-4.0 license.
A text dataset presents direct observations of host-guest complexation dynamics at solid-liquid interfaces. The data likely contains descriptions of frequency modulation and high-speed atomic force microscopy (AFM) results showing cooperative binding in densely assembled pillar[5]arene host structures. The dataset was authored by Hitoshi Asakawa and last updated on May 20, 2026.
32.8 KB of text data from figshare, authored by Hitoshi Asakawa and last updated in May 2026. The dataset describes the direct observation of association-dissociation dynamics in densely assembled pillar[5]arene host structures using frequency modulation and high-speed atomic force microscopy. It reveals positive cooperativity in guest binding at solid-liquid interfaces.
Spider DPO 1040 is a compact training dataset containing 1,040 preference pairs for Direct Preference Optimization, derived from frontier-model disagreements on the Spider V1 benchmark. It also includes 7,000 supervised training examples from Spider formatted for use with LLaMA-Factory. The dataset was created by jk200201 and was last updated on July 5, 2026.
Yi-Hui Zhou presents a hierarchical Bayesian framework for harmonizing collision cross section (CCS) measurements across ion mobility platforms. The dataset includes 840 measurements for 347 compounds from a multilaboratory study, used to validate the framework. The approach reduced intertechnology CCS variability by approximately 95% and lowered median absolute percentage error from 8.9% to 3.2%.