Loading...
Loading...
DNA/RNA sequences, gene expression, protein structures, metagenomics, single-cell sequencing
27,749 datasets
374.3 KB of differentially expressed genes identified via RNA sequencing. Norihiko Suzaki published this dataset under a CC-BY-4.0 license on figshare, with a last update in April 2026. The data compares gene expression between siSETD8-treated and control-treated SNU475 hepatocellular carcinoma cells.
Internal Consistency Reliability Coefficients for the AIMSP-C Final Model, authored by Shaoshen Wang. The dataset is a 5.5 KB Excel file last updated on April 13, 2026 and shared under a CC-BY-4.0 license.
Factor loadings for the AIMSP-C instrument, detailing results from initial and final statistical models. The dataset was authored by Shaoshen Wang and last updated on April 13, 2026. It is a 5.5 KB Excel file available under a CC-BY-4.0 license.
Ahmad Dweikat published a psychometric dataset containing item loadings, means, standard deviations, Cronbach’s alpha, composite reliability, and average variance extracted for an initial measurement model. The dataset is a single 9.5 KB Excel file, last updated in April 2026.
New York City shooting incidents recorded by the NYPD since 2006. The dataset includes location descriptions, precincts, and geographic coordinates for each event. It is part of a broader data collection on shootings, victims, and offenders hosted on the NYC open data portal.
The Texas Department of Insurance publishes a quarterly report of employers with active workers' compensation insurance coverage, known as subscribers. Data originates from insurance carriers via the IAIABC Proof of Coverage (POC) Release 2.1 standard, collected by the National Council on Workers’ Compensation Insurance (NCCI). This dataset includes employer names, addresses, policy dates, industry codes, and coverage provider information.
3,548 milk fat globule membrane proteins identified across rabbit colostrum and mature milk. The dataset includes 480 differentially expressed proteins and functional annotations from GO and KEGG analyses. Liangde Kuang published the data on figshare in March 2026.
A 2026 proteomics dataset from figshare profiles milk fat globule membrane proteins in rabbit colostrum and mature milk. The dataset contains 3,548 identified proteins, with 480 differentially expressed between the two milk stages. It was created by Liangde Kuang using data-independent acquisition quantitative proteomics.
Kaggle hosts this dataset of word embeddings derived from Bengali Wikipedia text. The dataset likely contains vector representations for words in the Bengali language, intended for natural language processing tasks. Its specific size, creation method, and update history are not detailed in the provided metadata.
Action Genome is a dataset for video understanding and action recognition. It is hosted on Kaggle, but detailed metadata such as author, organization, and creation date are not provided. The dataset's specific size, format, and annotation schema require verification after download.
A palynological report examines twenty-two samples from the S.P.L. No.1 (Birkhead) well. The analysis suggests the Permian-Triassic boundary lies between 3600 and 3700 feet and indicates a marine influence in the Upper Permian. The report was published by the Australian Ocean Data Network and was last updated in April 2026.
Laboratory experiment data from 2011-2013 measures the impact of simulated ocean acidification on the immune cells (hemocytes) of Alaskan Tanner crabs (Chionoecetes bairdi). Flow cytometry was used to analyze cell death, phagocytosis, and intracellular pH across three pH treatments (8.06, 7.80, 7.51) over two years. Results indicate increased hemocyte death and altered intracellular pH regulation under acidified conditions, suggesting potential impacts on crab immune health.
Bathymetry data for the Popes Eye area was collected by Deakin University Marine Mapping lab over two days in January 2018. The survey was conducted from the Motor Vessel Yolla using a Kongsberg EM2040c system as part of a Parks Victoria project to map marine parks in Victorian state waters. The data is not intended for navigational purposes.
Construction-Related Incidents is a dataset of construction-related incidents recorded by the New York City Department of Buildings (DOB). It contains records from the DOB Incident Database, with details on incident type, location, and outcomes. The dataset was last updated on April 3, 2026.
A 2021 bathymetric survey of Port Phillip Bay, Australia, conducted on December 13th by Deakin University's Marine Mapping lab. Data was collected from the Motor Vessel Yolla using a Kongsberg EM2040c sonar system to assess sediment movement over time. The dataset is provided by the Australian Ocean Data Network.
Survey and experimental data from two studies on Israeli public opinion during the 2023–25 conflict. The data was collected by Shingo Hamanaka and published via Harvard Dataverse in April 2026. It examines how perceptions of humanitarian risk, particularly regarding hostages and non-combatants, influence support for military escalation.
152.8 MB of genomic data files for the insect Sitotroga cerealella, including genome assembly, repeat annotation, gene annotation, and functional annotation. The dataset was authored by Zhichao Yan and last updated on April 9, 2026.
A three-day bathymetry survey of Port Phillip Bay was conducted in January 2018 by Deakin University. The data was collected using a Kongsberg EM2040c sonar system aboard the Motor Vessel Yolla. This work was part of a collaborative program with the University of Melbourne to map drift algae locations.
A bathymetry survey of the Portsea Hole in Port Phillip Bay, acquired over two days in January 2018. The data was collected by Deakin University Marine Mapping lab using a Kongsberg EM2040c sonar system onboard the Motor Vessel Yolla. This survey was part of a Parks Victoria project to map marine parks within Victorian state waters.
Over 915 million unique tweets, including retweets, were collected from the Twitter Stream, with dedicated gathering starting March 11th yielding over 4 million tweets daily. The dataset includes a cleaned version without retweets, covers multiple languages with higher prevalence of English, Spanish, and French, and provides pre-processed n-gram frequencies for NLP tasks. It offers longitudinal coverage from at least January 27th, with specific additions like Russian language tweets collected between January 1st and May 8th.