Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,694 datasets
Two Geoscience Australia surveys in 2003 and 2005 discovered submerged coral reefs in the Gulf of Carpentaria. The reefs were identified using new multibeam sonar technology and their age was determined via Uranium/Thorium dating of drill-core samples at the Australian National University. This discovery added a new coral reef province to Australia's marine zone.
Francisco J. Mendoza's research dataset documents the first detection of the parasite Theileria haneyi in horses in Spain. It contains results from analyzing blood samples from 222 equids across 19 Spanish provinces, with 3 positive for T. haneyi and 76 for T. equi. The dataset was last updated on June 4, 2026.
Supplementary data supporting a multi-cohort study on diabetic retinal neurodegeneration (DRN). The data likely contains proteomic and clinical variables used to develop a machine learning-based prediction model (Pro-DRN). It was authored by Huangdong Li and last updated on June 2, 2026.
Kingston City Council provides geospatial data on designated dog areas. The dataset defines zones where dogs are allowed off-leash, must be on-leash, or are prohibited, including exceptions for seasonal time restrictions. It was created by the GIS Team at City of Kingston and was last updated on 2026-07-16.
England's modelled annual live birth estimates for local authority districts, regions, and ITL2 areas. The Greater London Authority Demography team produces these estimates using GP registration counts to provide more timely data than official ONS figures, which have a 9-12 month publication lag. The dataset includes official ONS estimates from July 1992 to January 2025, interpolated monthly figures, and predicted births up to October 2025 with 95% prediction intervals.
Amino acid sequences and phylogenetic analyses for vertebrate oxytocin receptor (OTR) and vasopressin receptor (VPR) genes. The data includes sequences from 14 vertebrate species such as human, mouse, chicken, and zebrafish, curated from the Ensembl and Pre Ensembl genome browsers. The dataset was created by Daniel Ocampo Daza and is based on a 2012 study proposing an update to VPR gene nomenclature.
5.64 million rows of historical dock usage data for Melbourne's Bike Share program, which ended in November 2019. The dataset, provided by the City of Melbourne Open Data, includes station locations, bike counts, and empty slot counts. Readings were collected from sensor-equipped bike share pods between 2011 and 2017.
Western Australia's Carnarvon Shelf was surveyed in August and September 2008 by the CERF Marine Biodiversity Hub. The collaborative effort between the Australian Institute of Marine Science and Geoscience Australia collected physical and biological data across three main study areas. The report describes survey methods, collected datasets, and provides initial interpretations of the data.
An explanatory dictionary provides detailed definitions for about 68 key toxicological terms selected from the IUPAC 'Glossary of Terms Used in Toxicokinetics'. Monica Nordberg from Karolinska Institutet authored this resource to bridge linguistic and disciplinary gaps in understanding. The dictionary organizes terms under 22 main headings, aiming to clarify the relationship between chemistry and toxicology.
Monica Nordberg of Karolinska Institutet compiled this explanatory dictionary to clarify complex toxicological terms. It consists of about 68 terms selected from the IUPAC 'Glossary of Terms Used in Toxicokinetics', organized under 22 main headings. The dictionary provides full conceptual descriptions to overcome linguistic and interdisciplinary barriers in understanding.
The Canadian Importers Database (CID) provides summary reports and lists of companies importing goods into Canada for 2018. It is published by Innovation, Science and Economic Development Canada under the OGL-CA-2.0 license. The data is available in Excel formats (XLSX, XLS) and was last updated in July 2026.
2,638 papers from five geoscience journals were systematically surveyed for their use of color visualizations. The dataset classifies papers by color encoding issues, including rainbow color maps and red-green deficiencies, based on publications from 2005, 2010, 2015, and 2020. Richard Westaway of the University of Bristol compiled this data, which also incorporates a prior survey from Stoelzle and Stein (2021).
Multiple single-cell RNA sequencing datasets compiled by Jolene S. Ranek at the University of North Carolina at Chapel Hill for comparing temporal gene expression integration methods. The collection includes preprocessed data objects and raw files from public repositories like ArrayExpress and Gene Expression Omnibus. It covers biological processes such as mouse embryonic cell cycles, hematopoiesis, immune cell differentiation, and disease states including acute myeloid leukemia and multiple sclerosis.
541 mothers who gave birth in the last 12 months provided data on mental health and sociodemographics. The dataset supports validation of the French City Birth Trauma Scale for assessing posttraumatic stress disorder following childbirth. Vania Sandoz and colleagues at the University of Lausanne created this dataset for a 2022 study published in Psychological Trauma.
A seven-month community-based cross-sectional study from September 2024 to March 2025 surveyed adults in Kandahar province. The dataset contains knowledge, attitude, and prevention practice data for cutaneous leishmaniasis from 2,118 participants. It was authored by Bilal Ahmad Rahimi and published on figshare.
The DCASE 2018 Task 5 development dataset is a derivative of the SINS database, containing audio for detecting daily activities in a home. It consists of 72,984 ten-second audio segments from four microphone arrays, totaling approximately 200 hours of data. The segments are manually annotated for activities like cooking, dishwashing, eating, and social activity.
Retention scores for Internal Eliminated Sequences (IESs) and Optional Eliminated Sequences (OESs) from a study of natural eukaryotic genome editing. The dataset includes supplementary files with Spearman's correlation statistics for IESs of different lengths and for OESs. Data was produced by Estienne C. Swart at the University of Bern for the associated manuscript.
Elastic Net prediction models and LD reference data from the GTEx v8 project, supporting PrediXcan and MultiXcan analyses. The data was created by Alvaro Barbeira at the University of Chicago and is associated with a specific publication. Users must cite the source publication when using this data.
Rob Edinburgh from the University of Bath conducted a study assessing acute and chronic effects of nutrient-exercise timing on lipid metabolism and insulin sensitivity in overweight and obese men. The project comprised an acute metabolic study and a 6-week randomized controlled training study, involving men with a mean age of 30-35 years and a BMI of approximately 30-31 kg/m². Data likely includes measurements of lipid utilization, muscle fiber responses, oral glucose insulin sensitivity (OGIS index), and skeletal muscle adaptations.
Databases for the MyCodentifier pipeline, a bioinformatics tool for identifying Mycobacterium tuberculosis complex and Nontuberculous mycobacteria species from sequencing data. The collection includes seven major mycobacterial reference genomes, a Centrifuge classification database, and an hsp65 gene database. It was developed by researchers at Radboud University Nijmegen.