Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,705 datasets
AMBER Atlas of Instream Barriers in Europe is a geo-referenced database of 629,955 unique artificial instream barriers across 36 European countries. It was produced by the AMBER Project, funded under the European Union’s Horizon 2020 program. The dataset includes data for ground-truthing, barrier density modeling, and analysis of barrier under-reporting, as published in Nature.
Confidential personnel records for all management employees of a medium-sized U.S. service firm from 1969 to 1988. The data includes employee ID, age, sex, race, education, job title, salary, bonus, salary grade, and performance rating. Principal investigators from Harvard University Press used these records to study the firm's internal labor market and wage policy.
The Australian Ocean Data Network provides interpretations of U-Pb detrital zircon dating from offshore petroleum well cuttings. This analysis offers new information on sediment origin and changes in provenance for the Triassic Upper and Lower Keraudren deltas. The study aims to clarify potential Australian and non-Australian sediment sources for reservoir units in the Roebuck Basin on Australia's North West Shelf.
Computational workflows describe multi-step methods for data collection, preparation, analytics, and simulation that lead to new data products. These workflows inherently contribute to FAIR data principles by processing data with established metadata, creating metadata during processing, and tracking data provenance. The dataset, authored by Carole Goble of the University of Manchester, includes an example workflow for detecting variants in genome sequences specified in the Common Workflow Language.
Recordings of a wooden cube repeatedly dropped by a Kinova Gen3 Robot Arm. The dataset includes audio from a 7-channel microphone array and video from wrist and ceiling cameras, capturing the cube's impact sound and trajectory. It was created by Fanjun Bu of Cornell University to address robot error recovery based on object permanence.
53 residential building cohorts from Finland provide material volumes, masses, and intensities aggregated at four hierarchical levels. Data is based on digitized construction documents from Vantaa and Tampere archives, covering buildings from the 1940s to 2010s. The dataset was created by Tapio Kaasalainen of Tampere University.
1988 single-copy loci from 27 Caenorhabditis species form a supermatrix for phylogenomic analysis. The dataset includes multiple species tree files generated by ASTRAL-III, RAxML, and PhyloBayes, along with orthology clustering and gene structure statistics. Lewis Stevens from the University of Edinburgh created this dataset for comparative genomics research.
Supplementary data for 'Diversity in citations to a single study' published in Quantitative Science Studies. The dataset contains bibliometric data from Web of Science for all papers citing a specific cohort study up to 1984, including citation contexts manually recovered from 343 full-text papers. Author Rhodri Leng from the University of Edinburgh prepared the data, which is cleaned and coded for network analysis.
More than 50,000 recipes were collected from the online If-This-Then-That (IFTTT) service to characterize how humans use Internet-of-Things devices. The dataset contains basic information for each rule, including the trigger and action events, as well as the number of users per recipe. This collection supports research into evolving user behaviors and the evaluation of algorithms for smart spaces.
A 2022 dataset by David Cortés-Ortuño, Karl Fabian, and Lennart V. de Groot from Utrecht University contains simulation outputs and analysis scripts for mapping magnetic signals. The dataset includes MERRILL simulation scripts, output files, Jupyter notebooks for analysis, and pre-computed data files to calculate inversions and produce figures. It supports the associated preprint published in the Earth and Space Science Open Archive.
A 2014 project by Deltares, University of Groningen, and Utrecht University for Rijkswaterstaat and the Cultural Heritage Agency produced this dataset. It includes digital maps and databases defining the project area, providing a physio-geographical and archaeological data foundation, and the final archaeological expectation maps. The final product consists of a time series of nine maps covering successive archaeological periods from hunter-gatherers to modern times.
10 healthy participants (5 female, 5 male, mean age 24.1) underwent a discriminant delay fear conditioning experiment with auditory conditioned stimuli. The dataset includes skin conductance response measurements, CS and US information, and post-experiment subjective ratings of CS likelihood and preference. Matthias Staib from the University of Zurich contributed this data.
X-ray crystal structure refinements for normal human transthyretin and the amyloidogenic Val-30-Met variant, refined to 1.7-A resolution. The dataset includes R-values and stereochemical parameter standard deviations for bond and angle distances. The structures were refined by B.C. Braden, with improvements over a prior 1971 structure.
RNA-seq data from 27 glioblastoma multiforme (GBM) samples, as published in the manuscript 'Topographic mapping of the glioblastoma proteome reveals a triple axis model of intra-tumoral heterogeneity' by Lam et al. The dataset is associated with research from the University of Toronto and is available under an Open Access license.
Australia's marine region is classified into seascapes using an iterative, unsupervised 'crisp' ISOClass method in ERMapper. The methodology combines biophysical properties with consistent relationships to benthic biota and includes a validation using a 'fuzzy' classification on shelf data around Tasmania. This report, from the Australian Ocean Data Network, details the datasets and methods used, with a last update noted as 2026-06-23.
OpenKnot RNA Pseudoknots – RFDpoly 3D Structures contains computationally generated RNA 3D structures resembling pseudoknot topologies from the Eterna OpenKnot competition. The dataset was produced by RFDpoly through base pair template-conditioned structure generation and is authored by Andrew Favor of the University of Washington. It includes structures from competition rounds 7a and 7b, filtered by F1-score similarity to target secondary structures.
The Resolve dataset from a publication by Kristina Handler of ETH Zurich. It contains fragment-sequencing data designed to unveil local tissue microenvironments at single-cell resolution. The associated code for analysis is available on GitHub, and other related data can be found in GEO under accession GSE216189.
A spatial dataset containing the external extent of all Indigenous Land Use Agreements (ILUA) within Western Australia that are registered or in notification with the National Native Title Tribunal. The data is provided by the Western Australian Land Information Authority (Landgate) and was last updated in June 2026. It offers a spatial depiction of ILUAs, which can be made outside of the Native Title process but may still be part of a Native Title Determination.
A retrospective analysis of 195 adult patients with relapsed/refractory large B-cell lymphoma who underwent Next-Generation Sequencing. The dataset, authored by Rui Liu and last updated in June 2026, compares survival outcomes for patients with and without TP53 mutations treated with either CAR-T cell therapy or chemotherapy-based approaches.
A clinical dataset of 195 adult patients with relapsed/refractory large B-cell lymphoma who underwent Next-Generation Sequencing. The data compares survival outcomes for patients with and without TP53 mutations, stratified by treatment with CD19-directed CAR-T cell therapy or non-CAR-T chemotherapy. The dataset was created by Rui Liu and last updated in June 2026.