MMBench-Video is a quantitative benchmark designed to rigorously evaluate Large Vision-Language Models' proficiency in video understanding. It incorporates approximately 600 web videos and was created by opencompass. The dataset was last updated on October 9, 2024.
Use Cases
- Benchmarking model performance on holistic video understanding based on the described multi-shot evaluation framework.
- Evaluating long-form video comprehension capabilities based on the dataset's focus on extended video content.
- Assessing multimodal reasoning across visual and textual modalities based on the benchmark's design for LVLMs.
Strengths
- Designed as a rigorous quantitative benchmark for a specific evaluation purpose.
- Focuses on long-form and multi-shot video understanding, a complex task.
- Contains approximately 600 web videos for evaluation.
Limitations
- Row count and dataset size are unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- opencompass
- Collection Method
- Likely curated from approximately 600 web videos for benchmark construction.
- Freshness
- Last updated 2024-10-09 09:30:22; freshness should be verified.