Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,150 datasets
4.7 GB of external data files required for the NF-EnzymeFinder pipeline to mine novel enzymes directly from metagenomic reads. The data, authored by R Prabakaran, is hosted on figshare and was last updated on April 20, 2026. It is shared under a CC-BY-4.0 license.
Job postings from the City of New York’s official careers site, including internal and external opportunities. The dataset includes 30 columns such as Civil Service Title, Agency, Salary Range From, Salary Range To, Minimum Qual Requirements, and Preferred Skills. It is hosted on the NYC Open Data portal and was last updated on 2026-03-31.
This dataset documents a study establishing Tobacco rattle virus-mediated gene editing in groundcherry (Physalis grisea). It reports somatic editing frequencies of 80–95% for the PDS gene and up to 73% for CLV1, with heritable edits recovered in progeny. The study involved Cas9-expressing plants, with all five PDS-targeted plants producing edited offspring.
This dataset documents the outcomes of a virus-induced gene editing (VIGE) experiment in groundcherry (Physalis grisea), targeting the PDS and CLV1 genes. It reports somatic editing frequencies of 80–95% for PDS and up to 73% for CLV1, with heritable edits observed in progeny.
This dataset documents a study establishing Tobacco rattle virus-mediated gene editing in groundcherry (Physalis grisea). It reports somatic editing frequencies of 80–95% for the PDS gene and up to 73% for CLV1, with heritable edits recovered in progeny. The study involved Cas9-expressing plants, with results including fully albino seedlings and plants with increased floral organ number.
Aggregating survey results from 52 individuals in Nakuru County, Kenya, on fluoride exposure and oral health. It examines knowledge, attitudes, and behaviors regarding geogenic fluoride in drinking water, where samples exceeded WHO acceptable levels. The data was collected by Lavender Awino Okore and published in 2026.
Replication data for a political science article published in the European Journal of Political Research. The dataset, authored by Dominik Schraff and hosted on Harvard Dataverse, was last updated on 2026-05-28. It likely contains survey and regional economic data to analyze public support for fiscal transfers between territories.
Transparency materials for the study 'Buying Better Journalism? Media Subsidies and Public-Interest Content' were authored by Marcel Garz and published via Harvard Dataverse. The dataset was last updated on June 9, 2026, suggesting it may contain supporting data, code, or documentation for the associated research.
NanoFold Public is the public train/validation portion of the nanoFold protein-folding benchmark. It packages a compact, fixed, auditable subset of OpenProteinSet/OpenFold-derived protein structure training data for fast iteration on data-efficient folding models. The dataset contains 10,000 training chains and 1,000 public validation chains, with each row representing one protein chain.
An SMS dataset for multi-class classification tasks in the Azerbaijani language, distinguishing between legitimate messages (ham), spam, and smishing (SMS phishing). The dataset was created by Vusal Shahbazov and published in the journal Problems of Information Technology in 2026. It was last updated on the Hugging Face platform on May 8, 2026.
28 female secondary school students in a Ghanaian district were interviewed about their lived experiences with irregular menstrual cycles. The dataset contains thematic findings on their perceived causes, physical and emotional challenges, and coping strategies. Row and column counts are not specified.
A replication package for the academic paper "Dollar Asset Holdings and Hedging Around the Globe." It contains pseudo-data and codes, authored by Amy Huber and hosted on the Review of Financial Studies Dataverse. The package was last updated on June 9, 2026.
The International Ice Patrol compiles monthly and annual counts of icebergs drifting south across the 48° N latitude line in the western Atlantic Ocean. This dataset provides a continuous record from 1900 to the present day, offering a long-term view of iceberg activity. It is maintained by the NSIDCV0 organization.
Atlas is a dataset of microbiome samples created by outpost-bio. Each row represents one biological sample, containing aligned lists of taxonomic rank strings and their matching relative abundances. Additional columns record data type, sequencing method, pipeline version, and study accession.
A dataset from the Geosteering World Cup 2021 containing 176 expert readings of the same known geological scenario. The data likely captures variations in expert interpretation and labeling. It was created for a 2021 competition focused on geosteering.
Geoscience Australia compiled a historical record and bibliography of geological investigations in the Great Barrier Reef and adjacent regions. The compilation documents work by Geoscience Australia, its predecessor organizations, and collaborators, primarily concerning petroleum exploration interest. This record was published externally and last updated in April 2026.
The CAR LEADEX mission measured bidirectional reflectance functions for four common arctic surfaces: snow-covered sea ice, melt season sea ice, snow-covered tundra, and tundra shortly after snowmelt. The dataset provides insights into the variability of albedo in the arctic and was collected by the National Aeronautics and Space Administration. The dataset was last updated on March 13, 2026.
Community art is a relatively new form of art for the dataset's originating municipality, though the term has been established in English-speaking countries for decades. The data originates from the Dutch Ministry of the Interior and Kingdom Relations and is licensed under CC-BY-NC-4.0. Interest in this art form in the Netherlands has grown since 2003, aligning with political focus on societal cohesion.
Oxygymnocypris stewartii genome data includes assembly and annotation files. The dataset is 876.4 MB and was published by One more under a CC-BY-4.0 license. It was last updated on April 23, 2026.
Residential units and attributes of Address points for the District of Columbia, created as part of the Master Address Repository (MAR). The dataset includes condominiums and apartments, with addresses typically placed on buildings. It is provided by the Office of the Chief Technology Officer (OCTO) and DC Department of Consumer and Regulatory Affairs.