Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,749 datasets
Over 915 million unique tweets, including retweets, were collected from the Twitter Stream, with dedicated gathering starting March 11th yielding over 4 million tweets daily. The dataset includes a cleaned version without retweets, covers multiple languages with higher prevalence of English, Spanish, and French, and provides pre-processed n-gram frequencies for NLP tasks. It offers longitudinal coverage from at least January 27th, with specific additions like Russian language tweets collected between January 1st and May 8th.
PubMed abstracts on bioinformatics topics, fetched via the NCBI Entrez API. The dataset covers publications from 1990 through 2026 and is curated by Dr. Yash M Gupta for language model fine-tuning and text analysis.
Solar PV payback by province, Spain (2025) contains financial metrics for residential solar installations across Spain. The dataset likely contains payback periods, real internal rates of return (IRR), and 25-year savings projections for each of Spain's 47 provinces. It was sourced from Kaggle, but the author, organization, and specific data collection method are unknown.
City of Melbourne Open Data provides a 2014 register of council-owned land and building assets valued over $2.5 million. The dataset excludes 145 additional properties not listed. It is available in multiple geospatial and tabular formats including SHP, JSON, CSV, and GEOJSON.
203.5 MB of source data for Density Functional Theory (DFT) calculations and Molecular Dynamics (MD) simulations, authored by Linjie Zhi and shared under a CC-BY-4.0 license. The data, last updated in April 2026, supports research into ternary solvation sheath reconfiguration for sustainable cryogenic Li-Cl2 batteries. The files are packaged in multi-part archive formats (ZIP, RAR, Z01-Z05).
Diana Rojas Guerrero published decontaminated amplicon sequence variant (ASV) and operational taxonomic unit (OTU) tables on April 3, 2026. The 10.0 MB dataset contains processed COI, 16S rRNA, and ITS2 marker gene data from insect microbiome studies. It includes relative abundance calculations and spike-in corrected tables.
704 IAMT protein sequences were aligned to construct a maximum likelihood phylogenetic tree. The analysis used the JTT+R8 substitution model and 1000 ultrafast bootstrap replicates for branch support. The dataset, created by Bahar Saadaie Jahromi, represents a curated collection of sequences from NCBI and OneKP databases.
Supplementary material from a bioinformatics and machine learning study on ischemic stroke. The dataset is an Excel file published on figshare by Peilu Wang under a CC-BY-4.0 license, last updated in April 2026. Its 45.9 KB size suggests a limited scope, likely containing gene lists or analysis results.
22.4 KB Excel file containing supplementary figures (S1-S4) for a study on allele-specific suppression of pathogenic bestrophin-1 transcripts using CRISPR/Cas9 genome editing. The dataset, authored by Andrea Milenkovic, was published on figshare under a CC-BY-4.0 license and last updated in April 2026. Its specific content likely relates to experimental results and visualizations supporting the main research findings.
Analysis code for constructing a macrophage map across tissues in fetal pigs. The dataset, authored by Chenliang Lai and last updated in April 2026, is a 42.1 KB ZIP file containing scripts for single-cell profiling, likely related to fetal growth restriction studies.
An application of the Finite Volume Community Ocean Model (FVCOM v4.1) was run from May 24 to June 27, 2019 in the Discovery Islands region of British Columbia. The dataset contains modeled and observed temperature and salinity profiles used in a 2024 Ocean Science publication by Laura Bianucci et al. from DFO Ocean Sciences Division.
Action Genome Frames Part 3 is a dataset hosted on Kaggle, likely related to video understanding and action recognition. The platform tags suggest it contains video frames annotated for action analysis, potentially forming part of a larger series. Its specific content, scale, and authorship are not detailed in the available metadata.
Kaggle hosts the Action Genome Frames dataset, which appears to be a collection of video frames for action recognition tasks. The dataset's title and platform tags suggest it is designed for computer vision research involving human activities. Specific details on its size, creation date, and authorship are not provided in the available metadata.
Action Genome Frames Part 4 is a dataset hosted on Kaggle, likely related to video understanding and action recognition. The title suggests it is part of a larger collection, possibly containing annotated video frames. Metadata is minimal; the specific content, size, and collection details require verification after download.
Action Genome Frames Part 2 is a dataset hosted on Kaggle. The dataset's title suggests it relates to video-based action recognition, likely containing frames annotated for action understanding. Specific details regarding its size, columns, and creation are not provided in the available metadata.
Action Genome Frames Part 5 is a dataset hosted on Kaggle. The title suggests it is part of a series related to action recognition, likely containing video frames or annotations. Metadata is minimal; the specific content, scale, and creation details require verification after download.
Action Genome Frames Part 6 is a dataset hosted on Kaggle, likely containing video frames for action recognition tasks. The dataset's specific content, size, and creation details are not provided in the available metadata. Its platform tags suggest a focus on computer vision and video analysis.
Transcriptional profiling data from a pot trial investigating the glutathione metabolism response of Festuca sinensis grass, with and without Epichloë sinensis endophyte infection, to sodium selenite (Na2SeO3) stress. The dataset was generated by Lianyu Zhou using high-throughput RNA sequencing and was last updated in March 2026. It captures gene expression in plant shoots and roots over a 3-day period following treatment with 0, 20, and 50 mg/L Na2SeO3.
High-throughput RNA sequencing data investigates the glutathione metabolism of Festuca sinensis grass infected with Epichloë sinensis endophyte under sodium selenite stress. The dataset, authored by Lianyu Zhou and last updated in March 2026, captures physiochemical and transcriptional profiles from plant shoots and roots over a 3-day period. Pot trials treated plants with 0, 20, and 50 mg/L Na2SeO3 to explore mechanisms of selenium tolerance enhancement.
Supplementary materials from the Rare Kidney Disease Conference held in Palm Springs in December 2024. The 22.1 MB collection includes video recordings and document files shared by the conference organizer, Loma Linda University. The materials were uploaded by a Karger publishing administrator in April 2026.