39,594 action segments annotated across 55 hours of first-person video recorded in 32 kitchens. The dataset features 125 verb classes and 352 noun classes mapped to 11.5 million egocentric video frames.
Use Cases
- Train action recognition models to predict 'verb' and 'noun' labels from egocentric video segments
- Develop temporal action localization algorithms using the start_timestamp and stop_timestamp boundaries
- Analyze participant-specific behavior patterns using the participant_id and video_id columns
Strengths
- 39,594 temporal segments with precise start_timestamp and stop_timestamp metadata
- Dual-labeling system featuring 'verb' and 'noun' columns for fine-grained action recognition
- Covers 32 distinct kitchen environments across 4 different cities