58,000 human-annotated question-answer pairs based on 5,800 long-form videos from the ActivityNet dataset. The content covers diverse human activities and includes specific question categories such as motion, spatial relationships, and temporal reasoning.
Use Cases
- Train multi-modal models to predict the answer string given a video_id and a question
- Benchmark temporal reasoning capabilities by evaluating performance on the 'temporal' question category
- Develop video-language alignment layers that map natural language queries to specific visual frames
- Fine-tune vision-language models for long-context video understanding and reasoning
Strengths
- 58,000 natural language question-answer pairs
- 5,800 unique videos sourced from the ActivityNet action recognition collection
- Average video duration of 120 seconds per sample
- Question taxonomy includes motion, spatial, temporal, and attribute-based queries