1,065 natural language symptom descriptions labeled across 22 distinct medical diagnosis categories. The dataset focuses on fine-grained single-domain classification using English language text.
Use Cases
- Train a text classification model to map input_text symptoms to specific output_text diagnoses
- Evaluate the performance of medical entity extraction on natural language symptom descriptions in input_text
- Develop a diagnostic suggestion tool that processes English text inputs to predict one of 22 medical conditions
Strengths
- 1,065 rows of labeled medical symptom data
- 22 unique diagnosis labels within the output_text field
- Natural language symptom descriptions stored in the input_text column