CARV is a diagnostic benchmark created by researchers from Pennsylvania State University to evaluate compositional analogical reasoning in multimodal large language models. It assesses whether models can compose transformation rules from multiple image pairs via logical set operations like union and intersection. The dataset was last updated on July 13, 2026.
Use Cases
- Benchmarking multimodal LLMs' ability to synthesize rules from atomic visual changes.
- Evaluating model performance on logical set operations (union, intersection) applied to visual transformations.
- Diagnosing specific weaknesses in compositional reasoning within multimodal AI systems.
Strengths
- Designed as a diagnostic benchmark for a specific, advanced reasoning task.
- Focuses on compositional rule synthesis via logical set operations.
- Authored by researchers from Pennsylvania State University.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and file formats are unknown, which may limit suitability assessment.
Provenance
- Source
- huggingface
- Freshness
- Last updated 2026-07-13 21:59:34; freshness should be verified.