Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,080 datasets
Geoscience Australia Data assessed sediment distribution in Keppel Bay, a macrotidal environment interfacing the Fitzroy River catchment with the Great Barrier Reef shelf. The study classified seabed sediments into five distinct classes using sediment sampling, acoustic seabed mapping, and statistical techniques. The dataset was last updated on 2026-04-30.
Novel high-throughput experimentation sampling strategies require only 25% of the experiments compared to full factorial designs. Vincent Porte authored this dataset, which applies frugal sampling to four challenging metal-catalyzed cross-coupling reactions. The dataset was last updated on April 10, 2026.
A mathematical analysis of forest structural dynamics by Jian Zhou of Peking University. It demonstrates the significant impact of tree growth variance on forest size structure predictions, resolving a mismatch between observation and prior theory. The work identifies an asymptotically power-law relationship between tree size and growth rate variance.
Daniel T. Webb published a dataset on figshare in 2026 describing first-in-class Activin receptor-like kinase 2 (ALK2) degraders. The data includes information on compounds M4K3233 and M4K3250, developed as chemical tools for studying ALK2 degradation in diseases like fibrodysplasia ossificans progressiva and glioblastoma. The dataset is 1.6 KB in size and is available in CSV format.
18 spatial interpolation methods were compared for predicting seabed sand content within the Australian Exclusive Economic Zone (AEEZ). The study, using samples from Geoscience Australia's Marine Samples Database extracted in August 2010, found RFIDS and RFOK were among the most accurate methods across three tested regions. Model averaging further improved prediction accuracy, with the most accurate methods reducing error by up to 7%.
PUM-MATH contains pairwise prefix preference examples for gain-based evaluation of LLM reasoning. The dataset is associated with the paper 'From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning' and was authored by zhiqix. It was last updated on June 8, 2026.
Richard Goodman of IAOM published this dataset on July 11, 2026. The data relates to generating and checking polynomial receipts for Boolean satisfiability (SAT) and unsatisfiability (UNSAT) problems. It appears to involve a proof kernel for verifying computational results in formal logic.
The dataset 'The Arithmetic a Geometry Leaves Behind: Integral Fourier Duality on Abelian Varieties, Made Explicit' is authored by Richard Goodman from IAOM. It was last updated on July 11, 2026. The specific data format, size, and row count are not detailed in the available metadata.
Richard Goodman of IAOM authored this dataset on July 11, 2026. It provides a point-free and choice-free representation of Boolean and swap algebras, implementing concepts from Stone duality. The dataset's structure and specific contents are described in the accompanying academic work.
A mathematical dataset by Richard Goodman of IAOM, last updated on July 11, 2026. The content relates to prop-erasure, the Lawvere fixed point of vacuity, and equivariantization cores for fixed-point conformal nets. The specific data format and volume are not detailed in the available metadata.
Richard Goodman of IAOM authored a dataset titled 'The Public World Is a Quotient: Proof-Carrying Objectivity for a Black-Hole Archive Hypothesis'. The dataset was last updated on July 11, 2026. Its description suggests it relates to mathematical or philosophical concepts concerning objectivity and archival theory.
A dataset by Richard Goodman of IAOM, last updated on 2026-07-11, focuses on the machine-checking of foundational theorems in electromagnetism. It involves verifying Heaviside's Telegraph, Step, and Poynting Theorems from a single mathematical move, as indicated by the title. The description suggests the data is likely related to formal proofs or verification scripts for these theorems.
A formalization of a GΓΆdel universe and its shared fixed-point property with proof, likely representing a machine-checked mathematical model. The dataset was authored by Richard Goodman of IAOM and was last updated on July 11, 2026. Its specific data structure and size are not detailed in the available metadata.
Richard Goodman of IAOM authored this dataset related to the paper 'The Crystal Is the Verifier: Deciding String-Diagram Relations in Kashiwara's Combinatorial Skeleton'. The dataset was last updated on July 11, 2026. Its specific contents likely pertain to combinatorial structures and verification in representation theory.
A dataset related to the Clusterpath estimator of the Gaussian Graphical Model (CGGM), a method for variable clustering in graphical models. The dataset, authored by D.J.W. Touw and last updated on 2026-04-09, includes files such as TXT, ODS, PDF, R, MD, CSV, GZ, RDATA, GITIGNORE, and RPROJ, totaling 101.0 MB. It supports a convex optimization approach that encourages block-structured precision and covariance matrices.
RealMath-Eval is a benchmark for evaluating large language model judges on authentic human mathematical reasoning. The dataset was created by RicharMd and is associated with the paper 'RealMath-Eval: The Evaluation Gap in Judging Human Mathematical Reasoning'. The dataset was last updated on Hugging Face on June 10, 2026.
GeoTASO instrument data from the NASA B200 aircraft provides nitrogen dioxide (NO2) and formaldehyde (HCHO) trace gas slant column measurements over South Korea. This dataset was collected during the May-June 2016 KORUS-AQ field campaign, a joint study by NASA and Korea's National Institute of Environmental Research. It features coordinated airborne sampling up to 8 km altitude, focusing on the Seoul Metropolitan Area and surrounding waters.
A simulation experiment compares statistical and mathematical techniques for interpolating seabed mud content. The study, conducted by Geoscience Australia and published via the Australian Ocean Data Network, evaluates methods like random forest and ordinary kriging using cross-validation metrics. It was last updated in April 2026.
A figshare dataset by Xue Yuan, last updated April 2026, containing molecular structure data for a series of colony-stimulating factor 1 receptor (CSF1R) inhibitors. The data, stored in PDB format, relates to compounds like C52, which were designed for treating acetaminophen-induced acute liver injury. The dataset size is 365.9 KB.
Yu Zhang published a dataset on figshare in April 2026 detailing the design and optimization of a novel series of pyridazinone-based MAT2A inhibitors. The data likely contains results from structure-activity relationship studies, including IC50 and GI50 values for lead compounds. The dataset is 3.0 KB in size and is available in CSV format.