200 high-resolution video sequences containing over 300,000 frames of synchronized RGB and depth data for object tracking. The dataset covers diverse indoor and outdoor environments with manual bounding box annotations provided for every frame.
Use Cases
- Develop multi-modal fusion algorithms for object tracking using the RGB and depth frame pairs
- Train depth-only tracking models to identify and follow objects using only the depth map channel
- Benchmark tracker performance in long-term scenarios using the sequence-level challenge labels
Strengths
- 200 video sequences featuring synchronized RGB and depth modalities
- 300,000+ frames annotated with precise object bounding boxes
- Coverage of 15 distinct tracking challenges including occlusion, fast motion, and scale variation