Sign in to view source links and access this dataset
Description
OmniCap-IF is a benchmark dataset created by NJU-LINK for evaluating instruction following in omni-modal video captioning. It contains 480 videos and 1,920 instruction samples spanning tasks like understanding, generation, retrieval, and communication. Each sample pairs a prompt with fine-grained format and content checklists for evaluating structural, temporal, visual, audio, and audio-visual constraints.
Use Cases
Benchmarking model performance on instruction-following tasks based on the described understanding, generation, retrieval, and communication-oriented captioning tasks.
Evaluating adherence to structural and content constraints in video captioning based on the fine-grained format and content checklists.
Training or fine-tuning multimodal models on a diverse set of video-based instructions spanning visual, audio, and temporal modalities.
Strengths
Contains 480 videos and 1,920 instruction samples, providing a structured testbed.
Covers four distinct task types: understanding, generation, retrieval, and communication-oriented captioning.
Each sample includes fine-grained evaluation checklists for multiple constraint types (structural, temporal, visual, audio, audio-visual).
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
NJU-LINK
Freshness
Last updated 2026-06-08 10:09:56; freshness should be verified.
License is unknown; users should verify terms before use.