Causal-VL is a multimodal dataset for visual question answering focused on causal and physical reasoning. It contains 3,062 questions organized into 4 main categories and 16 subcategories. The dataset was created by author 'haorentang' and was last updated on May 22, 2026.
Use Cases
- Train models for causal reasoning based on visual scenes and causal graphs.
- Benchmark AI systems on physics-based anticipation tasks like collision prediction and fluid flow.
- Evaluate visual question answering performance on perception tasks such as scene reconstruction and mechanics reasoning.
- Develop models that understand intention speculation from visual data.
Strengths
- Structured organization with 4 categories and 4 subcategories each, providing clear task delineation.
- Contains 3,062 questions, offering a substantial number of data points for training and evaluation.
- Multimodal data includes images, videos, and structured annotations with causal graphs.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- huggingface
- Freshness
- Last updated 2026-05-22 10:47:24; freshness should be verified.