Loading...
Loading...
Medical imaging (X-ray, CT, MRI), electronic health records, clinical trials, ECG/EEG, pathology
14,809 datasets
Patho-Bench provides data splits for evaluating patch and slide encoder foundation models on whole-slide images. The benchmark was developed by the Mahmood Lab at Harvard Medical School and Brigham and Women's Hospital, funded by NIH NIGMS R35GM138216. Specific row counts, column details, and dataset size are not provided in the input.
TCGA-UT contains between 100,000 and 1,000,000 histology images extracted from uniform tumor regions within The Cancer Genome Atlas (TCGA) Whole Slide Images. Created by dakomura and updated in April 2025, it provides a benchmarking framework with predefined train, validation, and test splits for cancer classification.
HRScene is a unified benchmark for high-resolution image understanding, incorporating 25 real-world and 2 synthetic datasets. The collection includes images with resolutions ranging from 1,024 × 1,024 to 35,503 × 26,627, covering 25 scenarios. It was collected and re-annotated by 10 graduate-level annotators and was last updated on April 28, 2025.
Approximately 6,000 U.S. hospitals reported weekly COVID-19 hospitalization data to the CDC's National Healthcare Safety Network (NHSN). The dataset contains county-level metrics on admissions, inpatient and ICU bed occupancy, and weekly changes, aggregated from hospital reports. This dataset is archived and was last updated on May 3, 2024, after reporting requirements ended.
Weekly county-level COVID-19 hospitalization data from approximately 6,000 U.S. hospitals reported to the CDC's National Healthcare Safety Network. Metrics include admissions, inpatient and ICU bed occupancy, and population-adjusted rates, aggregated from December 15, 2022, until reporting ended in May 2024. The dataset is archived and no longer updated after May 3, 2024.
Published in Cell Patterns (2022) by futianfan, this benchmark dataset provides structured data for predicting the probability of clinical trial approval. It was developed to standardize performance evaluation for deep learning models, specifically the Hierarchical Interaction Network (HINT), within the therapeutic drug development domain.
Data from 1999 and even years 2002-2010 from the Behavioral Risk Factor Surveillance System (BRFSS) on adult oral health indicators. The dataset contains national estimates represented by the median prevalence among 50 states and the District of Columbia, prepared from BRFSS public use data sets. It is provided by data.cdc.gov and includes columns for prevalence values, confidence intervals, geographic locations, and demographic breakdowns.
25 simulated doctor-patient tele-consultation recordings totaling 4.2 hours of speech in Nigerian-accented English. Each interaction is captured as a single .wav file representing a complete medical dialogue between a healthcare provider and a patient.
Nova Scotia's data portal provides geographic locations for all hospitals in the province by civic address. The dataset includes facility names, types, towns, counties, and geometry coordinates. It was last updated in April 2025.
Four management zones define the operational structure of Nova Scotia's provincial health authority. The dataset contains geographic boundaries and identifiers for each zone, published by data.novascotia.ca. It was last updated in April 2025.
Data.cdc.gov hosts a dataset comparing the performance of multiple serological tests for detecting SARS-CoV-2 antibodies. The evaluation includes one in-house ELISA, two commercial chemiluminescence assays, and a surrogate virus neutralization test, using a panel of PCR-confirmed patient sera. The dataset was last updated on February 23, 2025.
An open-source collection of question-and-answer pairs for training AI models in clinical reasoning. The dataset, created by mattwesney and last updated on April 16, 2025, is designed to cover nuances in symptom presentation, diagnostic criteria, and treatment rationales for both physical and mental health.
Weekly respiratory virus-related hospitalization data aggregated to national and state/territory levels from approximately 6,000 U.S. hospitals. The CDC's National Healthcare Safety Network (NHSN) collected this data from August 1, 2020, to April 30, 2024, under a mandate and voluntarily from May 1, 2024. Metrics include counts and percentages for admissions, occupancy, and capacity specific to COVID 19 and influenza.
Potentially Preventable Readmission (PPR) rates for New York hospitals reveal facility-level performance on a key quality indicator. The dataset provides observed, expected, and risk-adjusted PPR rates for all payer beneficiaries. It is published by the New York State Department of Health via health.data.ny.gov, with records beginning in 2009.
CDC's archived dataset provides weekly COVID-19 hospitalization metrics aggregated to national, state, and regional levels. The data, reported by approximately 6,000 hospitals to the National Healthcare Safety Network, includes counts for admissions, inpatient bed occupancy, and ICU capacity. This dataset is archived and was last updated on February 23, 2025, as mandatory reporting ended in May 2024.
Cardiac surgery performance data includes case counts, observed and expected mortality rates, and risk-adjusted mortality rates for individual surgeons and hospitals in New York State. The dataset is produced by health.data.ny.gov and contains records for patients discharged between 2008 and 2010, with subsequent years appended. Physician-level data is reported for surgeons performing 200 or more procedures over three years or at least one surgery in each year.
HuggingFace user marcuscedricridia provides a cleaned version of the Medical-R1-Distill-Data dataset, last updated on April 3, 2025. The dataset, originally containing 22,000 entries, has been processed into a ShareGPT format with a single 'conversations' column. It consists of text-based dialogues between human and AI (gpt) roles.
Approximately 6,000 U.S. hospitals reported daily COVID-19 hospitalization data to the CDC's National Healthcare Safety Network. The dataset includes aggregated counts and metrics for hospital admissions, inpatient and ICU bed occupancy, and capacity at national, state, and regional levels. Data collection was archived after May 3, 2024, as mandatory reporting requirements ended.
A collection of demonstration datasets for robot manipulation domains, released with the Robomimic framework. The datasets are intended for benchmarking tasks and developing offline robot learning algorithms. The repository was last updated on April 13, 2025.
A dataset for medical question answering tasks, published on HuggingFace by Starlord1010. The dataset was last updated on June 4, 2025. Its specific content and size are not detailed in the available metadata.