E.T. Instruct 164K is a large-scale instruction-tuning dataset for fine-grained event-level and time-sensitive video understanding. It contains 101,000 videos across diverse domains and 9 event-level understanding tasks with instruction-response pairs, with an average video length of around 146 seconds. The dataset was created by PolyU-ChenLab and last updated on Hugging Face in September 2024.
Use Cases
- Fine-tuning video-language models for event-level understanding based on the 9 designed tasks.
- Training models for time-sensitive video reasoning based on the described temporal focus.
- Developing instruction-following AI for video analysis based on the instruction-response pairs.
- Benchmarking video understanding models across diverse domains based on the dataset's scope.
Strengths
- Contains 101,000 meticulously collected videos.
- Includes 9 well-designed event-level understanding tasks.
- Average video length is around 146 seconds, providing substantial context.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- PolyU-ChenLab via Hugging Face.
- Collection Method
- Meticulously collected videos with designed instruction-response pairs.
- Time Range
- null
- Freshness
- Last updated 2024-09-27 19:21:18; freshness should be verified.
- Geography
- null