Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,080 datasets
407,806 unique compact and extended X-ray sources are cataloged in this release from the Chandra X-ray Observatory. The Chandra Source Catalog version 2.1.1, updated in October 2024, provides over 100 uniformly calibrated positional, spatial, photometric, spectral, and temporal properties for each source, derived from 1,304,376 individual observation detections through the end of 2021. It is produced by the Chandra X-ray Center at the Harvard-Smithsonian Center for Astrophysics.
A synthetic dataset from a physiologically-based pharmacokinetic (PBPK) modeling study by Hyunseo Park, published in March 2026. The research simulates tigecycline exposure in plasma, epithelial lining fluid (ELF), and major organs to optimize inhaled dosing for Mycobacterium abscessus pulmonary infections. It includes model-based predictions for efficacy and safety thresholds across multiple dosing scenarios.
A mathematical paper by Žarko Mijajlović of the University of Belgrade presents a method for representing the inverse function of the cosmological scale factor a(t) as an elliptic integral. The work uses algebraic dependencies between cosmological parameters to compute special events in the universe's evolution in a uniform way. The dataset likely contains derived parameters or computational results supporting the paper's theoretical framework.
Dandan Guo from Huazhong University of Science and Technology authored a paper on the exponential stabilization of the wave equation with acoustic boundary conditions. The work uses Lyapunov and Riemannian geometry methods and applies the main theorem to wave equations with memory type acoustic boundary conditions. The paper includes an example application.
A theoretical work by G. Jothilakshmi of Alagappa University proposes an algebraic framework for analyzing fractional singular systems. The paper introduces a modern class of linear fractional singular delay systems with two orders and a decomposition method for matrix regular pencils. It includes a procedure for computing the reachable set and control input, illustrated with examples.
Noureddine Bouteraa of Université Oran 1 Ahmed Ben Bella authored a paper studying mild solutions of a fractional partial differential equation disturbed by multiplicative white noise. The work employs techniques of semigroup theory, Hausdorff measure, and Darbo fixed point theorem. The dataset likely contains mathematical or simulation results supporting this analysis.
A mathematical paper by Inès Feki from the University of Sfax proposes new supplements to linear operator perturbation theory. The work involves a non-analytic perturbation with multiple parameters and applies the theory to a Gribov operator in Bargmann space.
30 randomized controlled trials (N = 2,124) were analyzed to evaluate the dose-response effects of exercise on inflammatory biomarkers in overweight and obese postmenopausal women. The dataset, created by Gang Huang and last updated in March 2026, contains meta-analysis results from a systematic search of five databases up to January 2026. It includes standardized mean differences for biomarkers like CRP and TNF-α, with meta-regression results for exercise volume, duration, and intensity.
A meta-analysis of 30 randomized controlled trials (N = 2,124) investigating the dose-response effects of exercise on inflammatory biomarkers in overweight and obese postmenopausal women. The study, authored by Gang Huang and last updated in March 2026, systematically searched five databases and used meta-regression to analyze relationships between exercise parameters and biomarkers like CRP and TNF-α. Results indicate structured exercise significantly reduces TNF-α and CRP, with a trend suggesting higher intensity is more effective.
550,000 reasoning traces were distilled from the KIMI-K2.5 language model on high-reasoning tasks. The collection includes 2 billion tokens and is distributed across coding (60%), science (15%), math (10%), computer science (5%), logical questions (5%), and creative writing (5%). It was created by ansulev and last updated on Hugging Face in April 2026.
A geochemical dataset from estuarine sediment samples in Broad Sound, Queensland, analyzed using Q-mode and R-mode factor analysis, discriminant analysis, and regression. The data was published by Geoscience Australia Data and was last updated in March 2026. It identifies processes controlling concentrations of P2O5, Cu, Pb, and Zn in supratidal and intertidal zones.
This dataset contains experimental data from the optimization of 4-aniline substituted pyrido[3,2-d]pyrimidine derivatives as dual Pim/Mnk kinase inhibitors. It includes synthesized compounds with measured IC50 values against Mnk1, Mnk2, and Pim1 kinases, along with associated antiproliferative effects, solubility, and pharmacokinetic profiles. The data supports the development of compound 2j, a novel inhibitor with demonstrated in vivo antileukemia activity in MOLM-13 xenograft models.
1,000 Chinese multiple-choice math problems form this dataset, each annotated with a gold answer and a detailed rationale. It was introduced in the paper 'Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models'. The dataset was authored by SallyTan and last updated on Hugging Face in April 2026.
A simulation experiment from Geoscience Australia compares statistical and mathematical techniques for spatial interpolation of seabed mud content. The study analyzes factors including regions, sample densities, and secondary variables like bathymetry and distance-to-coast to assess prediction accuracy using cross-validation metrics. It identifies a novel combined method, random forest and ordinary kriging (RKrf), as the most robust, achieving up to 17% lower relative mean absolute error than a control method.
Estimated numbers and percentages of prevalent TB cases across urban and rural areas in 26 study countries from 2000 to 2024. Seyed Alireza Mortazavi created this dataset using a Bayesian multivariate regression model to estimate incidence and case detection ratios. The data was last updated in April 2026.
Supplemental material for the 2025 SOUPS paper by Tang et al. contains two files: a list of Symposium on Usable Security and Privacy (SOUPS) papers considered in the study and a table detailing the statistical tests and associated statistics examined. The dataset supports analysis of methodological reporting and interpretation in the usable privacy and security research domain.
A global statistical summary of column aerosol optical depth at 555 nanometers and monthly aerosol type frequency, averaged daily. The National Aeronautics and Space Administration produced this data from the Multi-angle Imaging SpectroRadiometer (MISR) instrument, which uses nine cameras at different angles to achieve global coverage in nine days. Data collection for this product is complete.
234.9 KB of raw experimental data underpins the statistical analysis of tower crane nonlinear swing behavior under stochastic wind excitation. The dataset was authored by Yu Sun and published on figshare in April 2026. It contains the measurements used to model load coupling mechanisms and varying windward pressures on flexible tower jibs.
Final training setups for baseline models compiled by Nadine Francis and uploaded to figshare on 2026-04-09. The 5.5 KB Excel file retains hyperparameters from original papers and marks common defaults. Missing values for optimizer, learning rate, schedule, momentum, or weight decay are indicated.
AIJJC_data_2026_LLM-optimization is a dataset from Kaggle focused on large language model optimization. The dataset's title suggests it contains parameters, configurations, or performance metrics related to tuning LLMs. Specific details on size, columns, and authorship are not provided in the available metadata.