4,754 multimodal videos totaling 215.4 hours of footage categorized into 6 violence classes including fighting, shooting, and explosions. The dataset provides synchronized audio and visual signals to support weakly supervised violence detection in diverse real-world scenarios.
Use Cases
- Train multimodal fusion models using the audio and visual streams to detect violent events
- Develop weakly supervised learning algorithms that map video-level violence labels to specific temporal segments
- Benchmark anomaly detection performance across 6 specific violence categories like 'Shooting' and 'Riot'
- Evaluate the impact of auditory cues on violence classification accuracy compared to visual-only methods
Strengths
- Contains 4,754 total videos with a cumulative duration of 215.4 hours
- Includes 6 distinct violence categories: Fighting, Shooting, Riot, Explosion, Car Accident, and Snatching
- Provides synchronized audio and video streams for multimodal feature extraction
- Designed for weak supervision where only video-level labels are provided during training