Sign in to view source links and access this dataset
Description
A Chinese-language dataset of 20,000 synthetic medical consultation dialogues generated by the DeepSeek-R1 model. The dataset, created by JackGao and last updated on April 15, 2025, includes patient questions, AI-synthesized responses, and department labels such as neurology and cardiology.
Use Cases
Training department classification models based on the 'label' field for patient queries.
Fine-tuning medical question-answering models using the 'question' and 'synthesized_text' fields.
Analyzing the characteristics of AI-generated medical text versus real doctor replies ('medical_replay').
Benchmarking model performance on Chinese medical dialogue tasks across different specialties.
Strengths
Contains 20,000 entries, providing a substantial volume of synthetic data.
Includes department labels for specialties like neurology and cardiology, enabling supervised classification tasks.
Provides both patient questions and AI-generated responses, creating a full dialogue structure.
Limitations
The AI-generated responses are from a DeepSeek-R1 FP8 model and have not undergone strict factual verification.
Medical queries and department classifications lack fine-grained governance and may contain inaccuracies.
The description notes that some doctor replies were obtained but not validated for authenticity.
Provenance
Source
huggingface
Collection Method
Synthesized by the DeepSeek-R1 model, with some doctor replies obtained but not validated.
Time Range
null
Freshness
Last updated 2025-04 15 10:49:38.
Geography
null
The dataset creator advises users to exercise caution and carefully scrutinize the data due to potential factual inaccuracies.