30,000 triplets combine real clinical notes from PMC-Patients case studies, synthetic patient-doctor dialogues generated by GPT 3.5, and structured patient information. The dataset was created by AGBonnet and last updated on 2024-01-24.
Use Cases
- Train text generation models based on synthetic patient-doctor dialogues.
- Develop note summarization tools based on real clinical notes from PubMed Central.
- Build models that integrate structured patient information with unstructured clinical text.
- Create conversational AI agents for medical training or simulation based on the dialogue data.
- Study the relationship between clinical notes and structured patient records.
Strengths
- Contains 30,000 data triplets from multiple sources.
- Includes real clinical notes extracted from PubMed Central case studies.
- Augments real data with synthetic dialogues generated by GPT 3.5.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is known, but specific file formats, size, and license details are unknown.
- Data may reflect bias inherent to the PMC-Patients source and the GPT 3.5 generation process.
Provenance
- Source
- AGBonnet via Hugging Face.
- Collection Method
- Combines real notes from PMC-Patients, synthetic dialogues generated by GPT 3.5, and structured patient information.
- Freshness
- Last updated 2024-01-24 10:38:13; freshness should be verified.