Sign in to view source links and access this dataset
Description
omi-health's Medical Dialogue To Soap Summary dataset contains 10,000 synthetic dialogues between patients and clinicians, generated using GPT-4 based on PubMed Central case reports. Each dialogue is paired with a corresponding SOAP summary, also generated by GPT-4. The dataset is split into 9,250 training, 500 validation, and 250 test entries.
Use Cases
Training models for automated SOAP note generation based on synthetic medical dialogues.
Evaluating the quality of AI-generated clinical summaries against a synthetic benchmark.
Fine-tuning language models for medical conversation understanding and information extraction.
Studying the structure and content of SOAP summaries within a controlled, synthetic dataset.
Strengths
Contains 10,000 synthetic medical dialogue-summary pairs, providing a substantial corpus for model development.
Includes a predefined split of 9,250 training, 500 validation, and 250 test entries for structured evaluation.
Dialogues and summaries are generated using GPT-4, likely offering coherent and structured text for NLP tasks.
Limitations
Data is synthetic, not derived from real clinical encounters, which may limit realism and generalizability to real-world settings.
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known, but specific file formats and data structure details are unknown from the provided metadata.
Provenance
Source
Generated by omi-health using GPT-4, based on case reports from PubMed Central (PMC).
Collection Method
Synthetic generation via large language model (GPT-4).
Time Range
null
Freshness
Last updated 2024-08-01 21:22:29.
Geography
null
License is unknown; terms of use must be verified before application.