Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,655 datasets
A 2026 network meta-analysis by Yicheng Ling, aggregating data from 26 studies with 9,169 subjects, quantifies the relationship between serum resistin levels and coronary heart disease severity. The analysis includes standardized mean differences and SUCRA rankings for disease stages from healthy controls to acute myocardial infarction. The dataset is derived from a systematic review registered with PROSPERO.
Yorkshire Water customer meter data includes both actual and estimated water consumption readings for domestic properties. The dataset has been anonymised to comply with the Data Protection Act and covers a 5-year period from 2010 to 2015. It is provided by the Calderdale Metropolitan Borough Council under the OGL-UK-3.0 license.
Yorkshire Water provides monthly meter readings expressed as mean litres per day for a representative sample of customers paying by rateable value. The data covers 2,470 anonymized properties across postal districts in Yorkshire and is used to estimate water use between 2010 and 2015. Calderdale Metropolitan Borough Council published this dataset, which has been anonymized for Data Protection Act compliance.
Gryllus bimaculatus, the two-spotted cricket, is a key hemimetabolous model organism for developmental biology, neuroscience, and regeneration. This dataset provides a chromosome-scale nuclear genome assembly, mitochondrial genome sequence, and structural and functional annotation files for the white-eyed mutant strain. The assembly was generated using a hybrid strategy combining Oxford Nanopore, PacBio HiFi, Illumina, and Hi-C scaffolding, representing a major improvement over a previous 2021 draft.
Decommissioned in May 2022, this dataset contains hydrodynamic model data for the Great Barrier Reef at 4 km resolution. The Ensemble Optimal Interpolation system assimilates temperature data from moorings, gliders, and satellite SST within a 6-hour window. The Australian Ocean Data Network produced this first version of a reanalysis for the GBR4 grid.
Nearly £400 million in payments from UK higher education institutions to seven major academic publishers between 2010 and 2014. The data was compiled by Stuart Lawson from Freedom of Information requests submitted through whatdotheyknow.com. Updates were added throughout late 2014 and early 2015, with notes on VAT treatment for electronic versus print subscriptions.
Environment Agency's Flood Map for Planning includes Surface Water Spatial Planning Extents. These datasets show land at risk from surface water flooding for three annual exceedance probability scenarios: 0.1%, 1%, and 3.3%. The data is designed for area-level risk indication, not individual property assessment.
Magnetopause crossing data from the THEMIS satellite probes between 2007 and 2016, classified for the manuscript by Staples et al., 2020. The data includes the time and location of each crossing event. The original THEMIS data used for classification is publicly available from NASA's CDAWeb.
Alphafold2 prediction output data for three distinct collections of fungal proteins. Dataset A contains predictions for 753 secreted proteins from Rhizophagus irregularis DAOM197198, Dataset B covers 10 fungal effectors, and Dataset C includes predictions for 454 matches from a MycFOLD-HMM search across Glomeromycotina genomes. The data was contributed by Albin Teulet of the University of Cambridge.
Brisbane City Council provides a dataset identifying over 2,000 parks across the city where specific sites can be booked for exclusive use. The dataset includes parks ranging from small local areas to large district parks, botanic gardens, and bushland reserves. It was last updated on 2026-06-28 and is available under a CC-BY-4.0 license.
A network-based toolbox for analyzing time series genome-wide Hi-C and RNA-seq data, developed by Stephen Lindsly at the University of Michigan. The toolbox provides methods to quantify network entropy, tensor entropy, and statistically significant changes in time series Hi-C data at different genomic scales. The data focuses on the 4D Nucleome, capturing genome organization and output over time.
Mexican American families provided data for Genetic Analysis Workshop 18, focusing on genes influencing complex traits. The dataset includes whole genome sequences for 464 key individuals from 20 families, SNP data for 959 individuals, and longitudinal blood pressure measurements over 20 years. Simulated phenotypes based on the real data, with 200 replicates for 849 individuals each, were created by Laura Almasy of the Texas Biomedical Research Institute.
Antarctic Circumnavigation Expedition data provides one-minute average horizontal wind velocity measurements from legs 0 to 4 of the 2016/2017 voyage. The data has been filtered for spurious observations and corrected for air-flow distortion caused by the ship's superstructure. Sebastian Landwehr from École Polytechnique Fédérale de Lausanne prepared the dataset, which is available under a CC BY 4.0 license.
Reference-guided genome assemblies and gene annotations for 16 species of Heliconius butterflies, created by Fernando Seixas of Harvard University. The dataset includes two scaffolded assembly versions per species aligned to H. melpomene or H. erato demophoon reference genomes, plus de novo mitogenome assemblies for all 16 species and an outgroup.
Lukasz Lebioda from the University of South Carolina determined the three-dimensional structure of yeast enolase using multiple isomorphous replacement and solvent flattening. The dataset describes a dimeric enzyme with an 8-fold beta+alpha-barrel domain exhibiting a novel beta beta alpha alpha (beta alpha)6 topology, refined to an R-factor of 17.0% with solvent molecules. The active site region and specific ligand residues like Asp246, Glu295, and Asp320 are detailed.
Two 2020 European projects' data collections enable comparison of different platform economy models and their connections to the UN Sustainable Development Goals. The dataset draws from the DECODE and PLUS projects, compiled by Mayo Fuster Morell of Harvard University. It presents the possibility to analyze the socioeconomic impact of platforms like Uber, Airbnb, and Deliveroo against alternative cooperative models.
The x-ray crystal structure of succinyl-CoA synthetase from Escherichia coli has been determined to a resolution of 2.5 Angstroms. The model has been refined to a conventional R factor of 21.6% with root mean square deviations from ideal stereochemistry of 0.022 A for bond lengths and 3.25 degrees for bond angles. The structure was determined by William T. Wolodko of the University of Alberta.
A 1.19-angstrom resolution X-ray crystallography model of the 131-residue rat intestinal fatty acid-binding protein without bound ligand. The refined model includes 237 solvent molecules, alternate conformers for 228 protein atoms, and details on 10 beta-strands and 2 alpha-helices. Giovanna Scapin from the Albert Einstein College of Medicine authored the associated research paper.
The World Trade Center Health Registry (WTCHR) monitors the health of 71,437 enrollees for 20 years. This analysis focuses on 8,418 adult survivors of collapsed or damaged buildings, with health outcome data collected via interviews from September 5, 2003, to November 20, 2004. The dataset was produced by the Agency for Toxic Substances and Disease Registry.
7243 unique reflections were used to refine the three-dimensional structure of recombinant human muscle fatty acid-binding protein to 2.1 Å resolution. The refined model has a crystallographic R factor of 19.5% and reveals a single bound fatty acid molecule within the protein's interior core. The structure was determined by Giuseppe Zanotti of the University of Padua using x-ray diffraction data.