1.27 million video-instruction pairs across categories such as video object grounding, tracking, and captioning. The content focuses on fine-grained object-level perception by mapping visual features to spatial-temporal coordinate tokens within video sequences.
Use Cases
- Train a grounding model using the <box> tokens and instruction strings to localize objects in motion
- Fine-tune Multimodal LLMs for temporal tracking using the frame-by-frame coordinate labels
- Develop spatial reasoning systems using the object-centric question and answer pairs to identify relationships between entities
Strengths
- 1.27 million video-instruction samples for fine-grained perception
- Includes <box> tokens representing normalized spatial-temporal coordinates [xmin, ymin, xmax, ymax]
- Covers diverse tasks including Video Object Grounding (VOG) and Video Object Tracking (VOT)