MMIR is a benchmark dataset designed to test multimodal large language models' ability to detect real-world cross-modal inconsistencies. It contains 534 carefully curated samples, each with a single semantic mismatch across five error categories, spanning webpages, slides, and posters. The dataset was created by rippleripple and last updated on 2025-02-25.
Use Cases
- Benchmarking model performance on cross-modal inconsistency detection based on the five defined error categories.
- Training models to identify semantic mismatches in multimodal documents like webpages and posters.
- Analyzing model failure modes in multimodal reasoning tasks using the curated samples.
Strengths
- Contains 534 carefully curated samples.
- Each sample injects a single semantic mismatch across five defined error categories.
- Spans diverse document layouts including webpages, slides, and posters.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Last updated 2025-02-25 03:30:54; freshness should be verified.
Provenance
- Source
- rippleripple on Hugging Face.
- Collection Method
- Curated for the Multimodal Inconsistency Reasoning (MMIR) benchmark.
- Freshness
- Last updated 2025-02-25 03:30:54.