Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
VSI-Super-Wild is a benchmark for evaluating multimodal models on spatial supersensing capabilities in long-form, in-the-wild videos. It was created by researchers from Tsinghua University, NVIDIA, and Stanford University for the ECCV 2026 conference. The dataset moves beyond short indoor clips and object-centric settings to study world state maintenance and prediction.
License is unknown, which may restrict usage. The full description is hosted externally, requiring a visit to the dataset page for complete details.