MSVD contains 1,970 videos, each paired with approximately 40 captions, providing a highly parallel corpus for language tasks. The dataset is split into 1,200 training, 100 validation, and 670 test videos, totaling over 80,000 captions. It was collected by David Chen and William B. Dolan for paraphrase evaluation research.
Use Cases
- Train video captioning models based on the multiple descriptive captions per video.
- Evaluate paraphrase generation systems based on the parallel captions for each video.
- Benchmark multimodal representation learning based on the aligned video-text pairs.
- Study linguistic diversity in visual descriptions based on the ~40 captions per video.
Strengths
- 1,970 videos provide a substantial visual corpus.
- Approximately 40 captions per video offer high linguistic parallelism.
- Official split provides 1,200 training, 100 validation, and 670 test videos for standardized evaluation.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Freshness should be verified; last metadata update was 2025-08-03.
Provenance
- Source
- Clone from 'friedrichor/MSVD' on Hugging Face, originally from research by David Chen and William B. Dolan.
- Collection Method
- Collected for paraphrase evaluation research.
- Freshness
- Last updated 2025-08-03 04:09:33