Video-R1 provides instruction-tuning data for reinforcing video reasoning in Multimodal Large Language Models. The repository includes a 165k example JSON file for supervised fine-tuning and a 260k example file for reinforcement learning training, sourced from datasets like CLEVRER and LLaVA-Video-178K. It was uploaded by the Video-R1 organization on April 11, 2025.
Use Cases
- Supervised fine-tuning of MLLMs based on the 165k chain-of-thought examples.
- Reinforcement learning training for video QA based on the 260k example dataset.
- Benchmarking model performance on video reasoning using the sourced datasets like NeXT-QA and PerceptionTest.
- Training models on multimodal tasks spanning video, image, chart, and spatial reasoning as indicated by the folder structure.
Strengths
- Includes two distinct, large-scale training sets: 165k examples for SFT and 260k examples for RL.
- Aggregates data from multiple established video and image QA benchmarks mentioned in the description.
- Provides a structured data format with fields like 'problem_id' and 'problem' as shown in the example snippet.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown for the constituent datasets, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- Video-R1 organization, via Hugging Face.
- Collection Method
- Aggregated and likely processed from multiple public video and image QA datasets.
- Time Range
- null
- Freshness
- Last updated 2025-04-11 09:09:04; freshness should be verified.
- Geography
- null