Loading...
Loading...
Mathematical datasets, statistical benchmarks, probability, optimization, operations research
3,047 datasets
A Bayesian model combines evidence from surveys, registers, and expert beliefs to estimate HIV infection prevalence. The analysis, by Christopher Jackson of the University of Cambridge, demonstrates Value of Information methods to identify which parameter uncertainties drive decision uncertainty and to prioritize future data collection. Supplementary materials for reproducing the work are available as an online supplement.
Bagged filter methodology combines an ensemble of Monte Carlo filters to address the curse of dimensionality in nonlinear, non-Gaussian systems. The research by Edward L. Ionides of the University of Michigan demonstrates its applicability to coupled population dynamics models, such as infectious disease transmission between cities. The method uses spatiotemporally localized weights to select successful filters at each unit and time.
A methodological paper by Jeffrey W. Miller of Harvard University introduces mixture of finite mixtures (MFM) models for Bayesian inference with an unknown number of components. The work includes real and simulated data, such as high-dimensional gene expression data used to discriminate cancer subtypes. Supplementary materials for the article are available online.
Jeffrey W. Miller from Harvard University introduces a novel Bayesian inference method that improves robustness to model misspecification. The approach conditions on the model generating data close to the observed data, approximated by tempering the likelihood. The paper illustrates the method with real and simulated data using mixture models and autoregressive models of unknown order.
A paper by Art B. Owen from Stanford University analyzes the statistical efficiency of thinning Markov chain Monte Carlo (MCMC) output. It provides theoretical examples and bounds, showing thinning can improve efficiency when computation cost per sample is high and autocorrelations decay slowly. Supplementary materials are available online.
Capture-24 is an activity tracker dataset for human activity recognition, superseded by a newer version. The dataset was created by researchers at the University of Oxford and is associated with studies published in Sage Journals and Nature Scientific Reports. It was used to test self-report time-use diaries against objective instruments and for statistical machine learning of sleep and physical activity phenotypes.
Australian nearshore regions are covered by daily gridded wind speed and direction data at 10m height, derived from Sentinel-1 Synthetic Aperture Radar (SAR) satellites. The dataset is produced by the Australian Ocean Data Network using the CMOD5N geophysical model and variational Bayesian inversion, with winds calibrated against a scatterometer database. Data are presented in delayed mode on a 0.01 x 0.01-degree regular grid.
A simulation experiment used samples from the Geoscience Australian Marine Samples database to compare statistical and mathematical techniques for predicting seabed mud content. Ten-fold cross-validation assessed prediction accuracy using metrics like mean absolute error and root mean square error. The study identified a novel combined method, random forest and ordinary kriging (RKrf), which reduced relative mean absolute error by up to 17% compared to a control.
An ecological observational study analyzing the spatial-temporal diffusion of sylvatic yellow fever cases during the 2017 epidemic in Espírito Santo, Brazil. The study found an incidence of 4.85 per 100,000 inhabitants and a case-fatality rate of 29.74%, with cases distributed across 34 of the state's 78 municipalities. The research was conducted by Priscila Carminati Siqueira from Universidade Federal do Espírito Santo using geostatistical analysis via ordinary kriging.
PMAF is a framework for designing, implementing, and proving the correctness of static analyses for probabilistic programs. The framework, developed by Di Wang at Carnegie Mellon University, handles challenging features like recursion, unstructured control-flow, divergence, nondeterminism, and continuous distributions. It has been used to reformulate existing intraprocedural analyses and implement a new interprocedural linear expectation-invariant analysis, with experiments on benchmark programs demonstrating its practicality.
Social Housing Statistics from the UK's Ministry of Housing, Communities and Local Government. This release was replaced by a new dataset on December 20, 2012. The data is designated as Official Statistics not designated as National Statistics.
Storm events on Australia's eastern and southern coast are analyzed using a 30-year observational timeseries. The dataset includes multivariate statistics for each event, such as maximum significant wave height, duration, and peak storm surge. This work is part of the Bushfire and Natural Hazards CRC Project Resilience to clustered disaster events on the coast.
Cycle 2 Water Framework Directive (WFD) unchanged class status at water body level for England in 2016. The dataset contains statistical confidence and derived certainty that a change in class has not occurred, assessed by detecting unchanged status against the 2015 baseline. It was created by the Environment Agency and has been retired, superseded by a newer record.
Cycle 2 Water Framework Directive class improvements at water body level for England in 2016, including statistical confidence and derived certainty of real improvement. The dataset was created by the Environment Agency and assesses upgrades in class status from 2015 to 2016.
Water Framework Directive Cycle 1 unchanged class status data for water bodies in England for the year 2016. The dataset, created by the Environment Agency, includes statistical confidence and derived certainty that a change in class did not occur compared to the 2009 baseline. This record has been retired and superseded by a newer version.
2016 data records Cycle 1 Water Framework Directive class deteriorations at the water body level for England only. The dataset assesses downgrades in class status from a 2009 baseline and includes statistical confidence and derived certainty metrics for each deterioration. It was created by the Environment Agency and is now retired, superseded by a newer record.
Cycle 1 Water Framework Directive (WFD) unchanged class status at water body level for England in 2015. The dataset contains statistical confidence and derived certainty estimates that a change in class has not occurred, assessed by comparing 2015 status against the 2009 baseline. It was created by the Environment Agency and is now retired, superseded by a newer record.
Environment Agency data contains Cycle 1 Water Framework Directive class improvements at the water body level for England in 2015. The dataset assesses upgrades in class status against a 2009 baseline and includes statistical confidence and derived certainty metrics for each improvement. This record has been retired and superseded by a newer version.
2015 water body status changes assessed against a 2009 baseline for England under the Water Framework Directive. The dataset includes statistical confidence and derived certainty metrics for each detected deterioration. It was created by the Environment Agency and is provided in XLSX format.
2014 data from England's Environment Agency assessing unchanged water quality status against a 2009 baseline. It contains statistical confidence and derived certainty that a change in class has not occurred for water bodies. This dataset has been retired and superseded by a newer record.