Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Med-HallMark is a benchmark dataset containing 750 image-question pairs for evaluating hallucinations in medical vision-language models. It includes three task types: conventional hallucination detection (499 pairs), counterfactual prompt-induced hallucination (111 pairs), and confidence weakening hallucination (140 pairs). The dataset was created by MM-Hallu and last updated on April 30, 2026.
The dataset page notes that an Image Report Generation (IRG) task requiring MIMIC-CXR/OpenI images is not included due to licensing restrictions.