Loading...
Loading...
Medical imaging (X-ray, CT, MRI), electronic health records, clinical trials, ECG/EEG, pathology
14,774 datasets
Centers for Disease Control and Prevention (CDC) data on Healthcare-Associated Infection (HAI) measures collected via the National Healthcare Safety Network (NHSN). The dataset provides information on infections related to devices like central lines and urinary catheters, or spread through patient contact. It is used to track and promote CDC-recommended infection control steps in hospitals.
National Institute of Standards and Technology provides a dataset of X-ray computed tomography (XCT) images for five additively manufactured cobalt chrome samples. Each sample contains approximately 1000 x 1000 x 1000 voxel three-dimensional images at a 2.5 ยตm resolution, with raw and segmented TIFF stacks. The samples were produced with varying laser powder bed fusion parameters to intentionally create defects.
HCUP Kids' Inpatient Database (KID) is the largest publicly available all-payer pediatric inpatient care database in the United States, containing data from 2 to 3 million hospital stays annually. Developed by the Agency for Healthcare Research and Quality, it is a sample of discharges from 4,000 U.S. hospitals for patients under 21 years old.
The Adverse Event Reporting System (AERS) database supports FDA post-marketing safety surveillance for all approved drug and therapeutic biologic products. It contains voluntary reports of adverse events and medication errors submitted by healthcare professionals, consumers, and manufacturers. The raw data files are not cumulative and require relational database expertise for analysis.
A collection of clinical trial records sourced from ClinicalTrials.gov, featuring structured metadata and eligibility information. The dataset includes pre-computed semantic embeddings for machine learning applications and was last updated by louisbrulenaudet on June 19, 2025.
Functioning as a multimodal collection of clinical data, pathology reports, slide images, molecular data, and radiology images for cancer patients, created by Lab-Rasool and updated in July 2025. It includes embeddings generated using models like GatorTron, MedGemma, Qwen, Llama, UNI, SeNMo, and REMEDIS. Specific row counts, column counts, and file formats are not provided.
1,200+ high-quality facial images of 400 individuals categorized by skin tone, skin type, and camera pose. The dataset provides detailed annotations for dermatological conditions including acne, rashes, and lesions across frontal and profile views.
Ukraine's municipal health center staffing information, sourced from the States site of Ukraine. The dataset describes the staffing of the Municipal non-profit enterprise "Hnivan Center for Primary Health Care" of the Hnivan City Council. It was last updated on July 16, 2025.
10,000+ MRI scans compiled by FLARE-MedFM for foundation model pretraining, as of July 2025. The dataset supports downstream tasks including liver tumor segmentation, cardiac tissue segmentation, and autism diagnosis. It is the official dataset for Task 4 of the MICCAI FLARE25 challenge.
A tool from the U.S. Department of Health & Human Services identifies addresses within designated health care shortage areas. It categorizes areas by specific needs, including Primary Care HPSA, Mental Health HPSA, Dental Care HPSA, or MUA/P. The data was last updated in July 2025.
A dataset hosted by makhresearch on Hugging Face, last updated in June 2025. It contains medical images for dermatology, annotated for segmentation and object detection tasks. The platform tags indicate the data includes bounding boxes and is formatted for YOLO, suggesting use in computer vision models.
Hugging Face hosts a dataset titled 'Medical Symptom Triage' uploaded by user sweatSmile. The dataset was last updated on August 10, 2025. Its specific content, size, and structure are not detailed in the available metadata.
The National Database for Clinical Trials Related to Mental Illness (NDCT) is an informatics platform for sharing de-identified human subject research data from NIMH-funded clinical trials. It contains data across biological and behavioral levels, including molecular, genetic, neural tissue, behavioral, social, and environmental interactions. The platform supports multiple data types such as text, numeric, image, and time series.
New York State (excluding New York City) surveillance data tracks pathogen prevalence in nymph deer ticks collected from May to September annually since 2008. The dataset includes counts of ticks tested and infected percentages for pathogens like Borrelia burgdorferi (Lyme disease), Anaplasma phagocytophilum, Babesia microti, and Borrelia miyamotoi. It is published by health.data.ny.gov.
10,000+ CT scans are provided for foundation model pretraining. The dataset supports downstream tasks including abdominal disease classification, abdominal lesion segmentation, abdominal organ segmentation, and lung lesion segmentation. It was created by FLARE-MedFM for the MICCAI FLARE25 challenge.
2008 onward surveillance data from collecting and testing adult deer ticks across New York State, excluding New York City, during the October to December season. The dataset includes pathogen infection percentages, tick population density, and totals for tested and collected ticks by county. It is published by the New York State Department of Health via health.data.ny.gov.
HCUPnet provides healthcare statistics for hospital inpatient stays and emergency department visits in the United States. The online tool allows users to generate tables and graphs on national and regional statistics and trends for community hospitals. Data originates from the Healthcare Cost and Utilization Project (HCUP).
Medical Symptom Triage Structured is a dataset uploaded to Hugging Face by user sweatSmile on August 10, 2025. The dataset's title suggests it contains structured information for medical symptom assessment and prioritization. Its specific content, scale, and structure require verification after download.
Mimic Cxr Az is a dataset of chest X-ray images hosted on HuggingFace. The dataset was uploaded by author nijatzeynalov and was last updated on August 11, 2025. Its specific content and scale are not detailed in the available metadata.
vidore's Syntheticdocqa Healthcare Industry Test dataset contains 1,000 pages sampled from PDFs collected from the internet. The dataset is designed to benchmark retrieval systems, specifically for medical documents, in realistic industrial applications. It was last updated on June 20, -2025.