WikiMed and PubMedDS: Large Datasets for Medical Concept Normalization
by Shikhar Vashishth / Carnegie Mellon University
Available on 1 platform
Sign in to view source links and access this dataset
Description
Two large-scale datasets contain over 58 million medical concept mentions linked to the Unified Medical Language System (UMLS). WikiMed, derived from Wikipedia, includes 1,067,083 mentions across 393,618 page texts, while PubMedDS, from PubMed abstracts, contains 57,943,354 mentions across 13,197,430 texts. Created by Shikhar Vashishth of Carnegie Mellon University, these datasets are intended for concept normalization research.
Use Cases
Training concept extraction models based on mentions with UMLS Concept Unique Identifiers (CUIs).
Evaluating concept normalization systems using the large-scale, automatically-annotated mentions.
Benchmarking against manually-annotated corpora like NCBI Disease Corpus, BioCDR, and MedMentions.
Studying the application of distant supervision from MeSH headers for biomedical text annotation.
Strengths
WikiMed annotations were manually evaluated on 100 random samples, showing 91% accuracy for UMLS CUIs and 95% accuracy for semantic type.
PubMedDS contains 57,943,354 mentions, providing a very large-scale resource for model training.
The datasets link mentions to a standardized vocabulary, with WikiMed covering 57,739 unique UMLS CUIs and PubMedDS covering 44,881.
Limitations
The description notes PubMedDS is not a comprehensive annotation of all medical concept mentions in abstracts, as only mentions located through distant supervision from MeSH headers were included.
Column-level documentation is absent; field semantics must be inferred after download.
Last update date is unknown; freshness unverified.
Provenance
Source
WikiMed derived from Wikipedia; PubMedDS derived from PubMed biomedical literature abstracts.
Collection Method
WikiMed created by crosswalking Wikipedia, Wikidata, Freebase, and NCBI Taxonomy to UMLS. PubMedDS mentions identified via distant supervision using MeSH headers and the scispaCy model.
PubMedDS is distributed as 30 individual files of approximately 1.5 million mentions each due to its size.