Loading...
Loading...
The dataset connects 20,000 videos to temporally annotated sentence descriptions. On average, each video contains 3.65 temporally localized sentences describing unique segments and multiple events.
The full description is hosted externally. The dataset is tagged with a non-standard license ('Licenseother').