Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,749 datasets
Analysis of a genome-wide Perturb-seq screen conducted on primary cells. The dataset was sourced from Kaggle, but details on the author, organization, and specific data volume are not provided. The description indicates a focus on functional genomics using single-cell RNA sequencing.
Overall ratings and condition scores from property inspections of New York City parks. Each record represents a single inspection, capturing details like Structural Condition, Safety Condition, and Cleanliness. The data is published by the City of New York and was last updated in March 2026.
Footprint outlines of buildings in New York City, sourced from data.cityofnewyork.us. The dataset includes 16 columns such as BIN, Construction Year, Height Roof, and Ground Elevation. It was last updated on March 29, 2026.
A schedule of public offerings and results for Oil Sands mineral rights leases in Alberta. The Government of Alberta administers these sales to companies for resource exploration and development. Data was last updated in April 2026.
Age-standardized rates of emergency department visits due to injury for Alberta and its health zones, expressed per 100,000 population. The data is provided by the Government of Alberta and was last updated in April 2026.
4,060 preprocessed microscopic images of 33 bacterial species, derived from the DIBaS dataset. The images were split and resized to 128x128 pixels and underwent blob normalization to reduce color bias. The dataset was created by Romain Mussard for the Meta-Album benchmark in March 2022.
840 remote sensing images of airplanes, resized to 128x128 pixels, form this dataset. It contains 21 different aircraft types, originally sourced from Google Earth satellite imagery and labeled by remote sensing specialists. The dataset was created by Philip Boser for the Meta-Album project in March 2022.
Molecular docking results for SBL virus protein structures include PDB coordinate files, PyMOL visualization sessions, and docking reports. The 46.7 MB collection was authored by Yao Xiong and last updated in April 2026.
The National Aeronautics and Space Administration project aims to fabricate silicon immersion gratings for infrared spectroscopy across 1150 to 6500 nm wavelengths. These devices are designed to support ground-based, airborne, and space-based spectrometers, offering a 3.44 times improvement in resolving power over conventional front-surface gratings of the same size. The dataset was last updated on March 13, 2026.
BIPA is a grapheme-to-phoneme dataset for Brazilian Portuguese, created by thiagomonteles. It provides transcriptions in the International Phonetic Alphabet and includes dialectal labels, derived from Wiktionary under CC BY-SA 4.0. The dataset features consistent symbol normalization and standardization of dialectal labels.
8.8 KB of tabular data from figshare, uploaded by Mikhail Sirenko on March 18, 2026, under a CC-BY-4.0 license. The dataset shows the robustness of key predictor effects to different operationalisations of responsibility perception, analyzing six structural measures. Self-efficacy and flood worry are highlighted as showing consistent significance across all operationalisations.
The Macreadie Seagrass bathymetry survey was acquired by Deakin University Marine onboard the Motor Vessel Yolla over two days in December 2015. Data collection used a Kongsberg EM2040c instrument. The dataset is hosted by Geoscience Australia Data.
Crayotter LongShOTBench Cases (23-case format) is a dataset hosted on Kaggle. The title suggests it is a benchmark for evaluating long-context reasoning in language models. Its specific content, size, and authorship are not detailed in the provided metadata.
Crayotter LongShOT Baseline Outputs – Batch 2 is a dataset published on Kaggle. The dataset likely contains benchmark results or predictions from a machine learning model, as suggested by its title and platform tags. Specific details regarding its size, columns, and creation are not provided in the available metadata.
Crayotter LongShOT Baseline Outputs – Batch 1 is a dataset published on Kaggle. The title suggests it contains outputs from a baseline model run, likely for a machine learning task. The dataset's specific content, size, and creation details are not provided in the available metadata.
Kaggle hosts this dataset of baseline outputs from the Crayotter LongShOT project. The dataset likely contains predictions or scores from a model run, intended for comparison and evaluation. Specific details such as the number of rows, columns, and the originating author are not provided in the available metadata.
Kaggle hosts baseline outputs from the Crayotter LongShOT project. The dataset likely contains predictions or scores from initial model runs intended for comparison. Its specific contents, size, and creation details require verification after download.
Kaggle hosts baseline outputs for the Crayotter LongShOT project. The dataset likely contains predictions or scores from benchmark models. Its specific content, size, and creation details require verification after download.
HUH7 RNAseq matrix normalized with EdgeR and Voom, associated with a phenotype file. Christophe Desterke published the 1.9 MB CSV file on figshare under a CC-BY-4.0 license. The dataset was last updated on 2026-04-10.
85,000 Python functions paired with short natural-language instructions derived from repository docstrings. The dataset was created by NickIBrody and last updated on April 20, 2026. It includes deterministic train, validation, and test splits.