Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,749 datasets
A 2020 dataset from OpenML for evaluating classification models on independent and identically distributed tabular data. It was curated by the TabArena team for their Tabular ML IID Study, focusing on features from a single 10 GHz frequency for detecting contaminants in food products. The original research was published by Ricci et al. and Urbinati et al. in IEEE journals and conferences.
North Atlantic Ocean data was collected by NOAA Ship OKEANOS EXPLORER during the Atlantic Seabed Mapping International Working Group (ASMIWG) Bermuda Mapping expedition from July 12 to July 31, 2018. The dataset includes shipboard sensor measurements for navigation, meteorology (conductivity, wind, temperature), and oceanography (bathythermograph, sound velocity probe, thermosalinograph), along with profile data from ASVP and XBT casts. It is managed by NOAA's National Centers for Environmental Information (NCEI).
Tomasz Szopiński's dataset contains survey data from a pilot study on the acceptance of AI-based tools for personal financial decision-making. The data was gathered via a CAWI survey from 371 respondents at three Polish universities. The study tests a dedicated scale based on the Technology Acceptance Model (TAM), examining constructs like perceived ease of use, usefulness, attitudes, intentions, and actual use.
166 transcriptomes from 75 metazoans were analyzed to investigate the relationship between transcript diversity and effective population size. Kai Mi authored this dataset, which was last updated on March 19, 2026. The analysis supports the hypothesis that transcript diversity is largely deleterious and declines with increasing effective population size.
Metagenomic and viromic data explores viral characteristics and virus-host interactions in sewer sediments from three distinct urban functional areas. The study includes a global prevalence analysis of key viral hosts across sewers in 76 cities from six countries. It investigates the dual role of viruses as metabolic tuners in sulfur dynamics and their potential for alleviating sewer corrosion.
A 2026 study by Ammar F. Ibrahim presents experimental and computational data on a nickel-catalyzed cross-dehydrogenative coupling method for synthesizing β,γ-unsaturated ketones. The dataset includes results from an array of substrate combinations and comparative analysis of oxidants like di-tert-butyl peroxide. Computational studies provide quantum chemical estimates of thermodynamics for the proposed double hydrogen-atom-transfer mechanism.
Experimental and computational chemistry data supports a nickel-catalyzed cross-dehydrogenative coupling process for synthesizing β,γ-unsaturated ketones. The dataset includes results from an array of substrate combinations and comparative analysis of oxidants like di-tert-butyl peroxide. Computational studies provide quantum chemical estimates of thermodynamics for the proposed double hydrogen-atom-transfer mechanism.
May 11-14, 2016 data collection by NOAA using a Riegl VQ880G sensor. The dataset covers approximately 85 square miles along the shores of Boca Grande, Florida, and contains 1,053 individual 500 m x 500 m lidar tiles. Data is provided in LAS 1.2 format with classifications for ground, water surface, bathymetric bottom, and other features.
St. Jeromes Creek, Maryland, is covered by approximately 34 square miles of topobathy lidar data collected on August 5, 2017. The National Oceanic and Atmospheric Administration acquired the data using a Riegl VQ880G sensor from an airplane. The dataset consists of 342 tiles, each 500 meters by 500 meters, with points classified for ground, water surface, and bathymetric bottom.
Multiple research cruises aboard the R/V Marcus G. Langseth collected surface underway measurements of carbon dioxide partial pressure, salinity, and sea surface temperature across global oceans from 2011 to 2015. Data from four distinct voyages cover the Arctic, Pacific, Atlantic, and Mediterranean seas, including several U.S. National Marine Sanctuaries. This dataset supports the study of ocean carbon uptake, acidification, and air-sea gas exchange dynamics.
Hydrographic data from 52 days in 1994 includes measurements of chlorofluorocarbons (CFC-11, CFC-12, CFC-113), dissolved oxygen, salinity, and water temperature from the South Pacific Ocean's Ross Sea. The dataset, collected via CTD and bottle instruments aboard the R/V Nathaniel B. Palmer, supports the CLIVAR program's goal of quantifying changes in ocean heat, freshwater, and carbon storage. It provides a discrete sample baseline for studying ocean circulation and anthropogenic tracer distribution in the Southern Ocean.
342 lidar tiles covering approximately 34 square miles of shoreline were collected by NOAA on April 9, 2017. The data was acquired using a Riegl VQ880G sensor and includes classifications for ground, water surface, bathymetric bottom, and other features. This point cloud data is provided in LAS 1.2 format.
Differentially expressed genes were identified in maize coleoptile libraries using RNA-seq analysis. The dataset, created by Yao Wang and last updated on April 13, 2026, is available as a 12.5 MB Excel file under a CC-BY-4.0 license. It compares gene expression between ZmSKIP overexpressing (OE2) and wild-type (WT) maize plants.
Boca Grande, Florida, is covered by approximately 85 square miles of topobathy lidar data collected by the National Oceanic and Atmospheric Administration from May 11 to 14, 2016. The dataset consists of 1,053 tiles, each 500 meters by 500 meters, in LAS 1.2 format. Data points are classified into categories including ground, water surface, bathymetric bottom, and noise.
9.5 KB Excel spreadsheet by Lumila Paula Menéndez, last updated in April 2026. This dataset provides a curated list of key variables for evaluating the integrity of ancient DNA samples. It likely contains metrics for assessing authenticity, quality, and contamination levels.
Entropy-ratio maps from the Australian Ocean Data Network enable the mapping of surficial coral reef facies at a user-chosen resolution. The method uses a ternary classification based on detritus, framework encrustation, and pavement, subdivided by their degree of mixing. An example application is provided for the Great Barrier Reef.
A 10.3 KB Excel file listing genome maintenance genes from the Mouse Genome Informatics (MGI) database. The dataset, authored by Mumingjiang Munisha and last updated on 2026-04 -13, specifically includes genes where perturbations have been documented to cause placental phenotypes.
A small dataset of International Classification of Disease codes for mosquito-transmitted arboviruses. The 5.5 KB Excel file was authored by Maria Elizabeth Mitri and last updated on April 13, 2026.
5.5 KB of parameter estimates from a sensitivity analysis comparing GLM Gamma regression models for maternal age at first birth, with and without socioeconomic variables. The dataset was authored by Adimias Wendimagegn Agegnehu and last updated on April 13, 2026. It is available under a CC-BY-4.0 license.
Parameter Estimates of Maternal Age at First Birth from Gamma Regression Models (GLM and GLMM) with Geographic, Residential and Religious Covariates. The dataset was authored by Adimias Wendimagegn Agegnehu and last updated on April 13, 2026. It is a 9.5 KB Excel file available under a CC-BY-4.0 license on figshare.