3.4 million video clips of 61 frames each, categorized into Cut (C), Transition (T), and Empty (E) classes. The data is sourced from AutoShot, ClipShots, and Pexels, featuring both real-world footage and synthetic transitions centered on the 31st frame.
Use Cases
- Train a binary classifier to distinguish between 'E' (Empty) and boundary frames ('C' or 'T')
- Develop temporal neural networks to detect gradual transitions using the 61-frame sequence context
- Benchmark shot boundary detection algorithms on the specific 'C' (Cut) and 'T' (Transition) labels
Strengths
- 3.4 million uniformly structured video clips containing exactly 61 frames each
- Labels applied specifically to the 31st frame to indicate boundary presence
- Three distinct classification labels: 'C' (Cut), 'T' (Transition), and 'E' (Empty)
- Aggregated from multiple sources including AutoShot, ClipShots, and crawled Pexels videos