267 medical image-caption pairs sampled at a 0.33% rate from the ROCOv2 radiology corpus. The dataset includes diverse imaging modalities such as CT, MRI, and X-ray paired with clinical descriptive text.
Use Cases
- Fine-tune vision-language models for medical image captioning using the image and caption fields
- Benchmark zero-shot image retrieval performance on a small-scale clinical subset
- Validate data loading and preprocessing pipelines before scaling to the full ROCOv2 corpus
Strengths
- Contains approximately 267 image-caption pairs sampled from the eltorio/ROCOv2-radiology repository
- Includes multimodal radiology images covering CT scans, ultrasound, and PET scans
- Features paired natural language captions describing anatomical structures and pathological findings