151,000 samples focused on video spatial reasoning designed to reinforce Multimodal Large Language Models (MLLMs). The dataset supports the SpaceR framework, providing data to improve model performance in understanding spatial relationships and dynamics within video sequences.
Use Cases
- Fine-tune Multimodal Large Language Models (MLLMs) using the 151,000 video spatial reasoning samples
- Benchmark the spatial logic of video-language models against the SpaceR dataset standards
- Develop reinforcement learning pipelines for MLLMs to improve accuracy in identifying spatial relationships in dynamic video
Strengths
- 151,000 samples dedicated to video-based spatial reasoning tasks
- Released under the CC BY-NC 4.0 license for non-commercial research
- Associated with the 2025 SpaceR research paper on reinforcing MLLMs