Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,749 datasets
Twenty-four red fox femorotibial joints reconstructed three times each for measurement variability analysis. The data consists of stacked DICOM files for each joint, produced by James Miles. Header metadata can be extracted using a provided ImageJ macro.
158 hybrid genome assemblies for 20 bacterial isolates compare PacBio and Oxford Nanopore long-read technologies combined with Illumina short reads. Assemblies were produced using the Unicycler tool with four distinct long-read preparation strategies: basic, corrected, filtered, and subsampled. The data supports a 2019 study by De Maio, Shaw, and the REHAB consortium on the accuracy of long-read sequencing methods for complex bacterial genomes.
The Marine Microbial Eukaryotic Transcriptome Sequencing Project (MMETSP) dataset contains cultured samples of pelagic and endosymbiotic marine eukaryotic species representing more than 40 phyla. It provides de novo transcriptome assemblies, peptide translations, gene annotations, and gene expression quantifications for these organisms. The data was generated using methods described in the Eel Pond khmer protocols and associated scripts are available on GitHub.
Tables with estimates of external costs of freight transport for each transport mode and EU Member State. These figures can be used as reference for the comparison of different measures to reduce externalities from transport. The dataset was authored by Panayotis Christidis of the European Commission.
Two Geoscience Australia surveys in 2003 and 2005 discovered previously unknown submerged coral reefs in the Gulf of Carpentaria. The reefs were identified using new multibeam sonar technology and their age was determined via Uranium/Thorium dating of drill-core samples at the Australian National University. This discovery established a new coral reef province within Australia's marine zone.
Two genome assemblies and associated sequences support a study on cryptic species delineation in marine ciliates. The data includes micronuclear genome assemblies, macronuclear transcriptome assemblies, and rRNA barcode sequences for Schmidingerella arcuata and Schmidingerella meunieri. It was published by Susan A. Smith et al. in Molecular Ecology Resources in 2022.
Late Quaternary pollen data from 2,831 records harmonized into a single chronology framework. The dataset includes age control points and metadata from the Neotoma Paleoecology Database and supplementary Asian sources. It was created by Chenzhi Li at the Alfred-Wegener-Institut Helmholtz-Zentrum für Polar- und Meeresforschung and is accompanied by R code for calculation and comparison.
Summary statistics from three genomic analyses for kidney disease research, created by Hongbo Liu. Kidney mQTL mapping was performed on 443 trans-ancestry individuals, identifying 139,313 mCpGs and 13,771,378 significant SNP-mCpG pairs. An eGFRcrea GWAS meta-analysis integrated five studies across 1,508,659 individuals, finding 90,950 genome-wide significant variants, while a kidney eQTL meta-analysis of 686 individuals identified 10,430 eGenes and 1,222,250 significant SNP-gene pairs.
Radboud University Nijmegen provides supplementary data for the scANANSE methodology paper, which enables gene regulatory network and motif analysis of single-cell clusters. The repository contains pre-processed Seurat and Scanpy objects derived from 10x Genomics Cell Ranger ARC data for sorted human PBMC and granulocyte cells. These objects are intended for running the scANANSE pipeline to infer transcription factor activity and regulatory networks from combined RNA and ATAC sequencing.
Parviz Ghaderi provides electrophysiological recordings from three transgenic mouse lines (VGAT-ChR2-YFP, SST-ChR2-YFP, PV-ChR2-YFP) under urethane anesthesia. The dataset includes raw and processed membrane potential data sampled at 20 kHz from parvalbumin-positive (PV), somatostatin-positive (SST) interneurons, and nearby pyramidal neurons in the primary visual cortex. Data originates from experiments approved by the RIKEN Brain Science Institute Animal Experiment Committee and is associated with a 2017 Scientific Reports publication.
A complete mitochondrial genome assembly for the cactus species Lophophora williamsii (peyote), spanning 2,422,778 base pairs. The dataset includes results from analyses of repeats, codon usage, RNA editing, gene transfer, phylogenetics, and selection pressure. It was authored by Xingliang Liu and uploaded to figshare on 2026-06-02.
The first complete mitochondrial genome of the Lophophora williamsii (peyote) cactus, native to the Chihuahuan Desert. Xingliang Liu assembled and annotated the 2,422,778 bp genome using PacBio HiFi long-read sequencing, with analyses including repeat identification and phylogenetic reconstruction. The dataset was last updated on June 2, 2026.
Historic de-identified medical data from five metropolitan areas over 23 months was used to evaluate outbreak detection algorithms. The dataset includes International Classification of Diseases, Ninth Revision (ICD-9) codes related to respiratory and gastrointestinal illness syndromes. David W. Siegrist of the Potomac Institute for Policy Studies led this effort to assess algorithm sensitivity and timeliness.
A clinical study of 38 pregnant women assessed the psychological and behavioral effects of false positive ultrasound soft marker screenings. The research, led by Sylvie Viaux‐Savelon of the Centre National de la Recherche Scientifique, measured anxiety, depression, maternal representations, and mother-infant interactions at multiple points from the third trimester to two months postpartum. Data likely includes psychological scale scores and behavioral coding from videotaped interactions.
2.2-Angstrom resolution crystal structure of recombinant calmodulin from Drosophila melanogaster, refined with a crystallographic R value of 0.197. The model includes 1,164 protein atoms, 4 calcium ions, and 78 water molecules, revealing structural differences from mammalian calmodulin. The data was produced by Denise Taylor of the Howard Hughes Medical Institute.
T.L. Poulos from the University of California, Irvine refined the crystal structure of the major lignin peroxidase isozyme from Phanerocheate chrysosporium. The model, refined to an R-factor of 0.15 for data between 8 and 2.03 Angstroms, includes 2 molecules per asymmetric unit, calcium ions, glucosamine modifications, and 476 water molecules. The refinement confirms earlier conclusions and details structural similarities and differences with cytochrome c peroxidase.
Drug-related deaths in Scotland are broken down by age, sex, Health Board, and Council areas. The data is designated as National Statistics and produced by National Records of Scotland. It was last updated on 2026-07-08.
Household estimates by local authority area in Scotland, including information on household type and numbers of vacant dwellings. The data is produced by National Records of Scotland and designated as National Statistics. The dataset was last updated on 2026-07-08.
Monthly figures on museum and gallery visits are provided by the Department for Digital, Culture, Media and Sport. The data is designated as Official Statistics not designated as National Statistics. The dataset was last updated on 2026-07-08.
Official statistics from National Records of Scotland list the names given to newly born children for a calendar year. The dataset includes the frequency of each name. It was last updated on 2026-07-08.