Babillage is a multimodal benchmark dataset introduced alongside MoshiVis. It contains three common vision-language benchmarks—COCO-Captions, OCR-VQA, and VQAv2—converted into spoken form for evaluating Vision Speech Models. The dataset was created by kyutai and last updated on March 21, 2025.
Use Cases
- Benchmarking Vision Speech Models based on spoken versions of COCO-Captions
- Evaluating model performance on spoken question-answer pairs based on OCR-VQA
- Testing multimodal understanding on conversational audio dialogues based on VQAv2
Strengths
- Based on three established vision-language benchmarks: COCO-Captions, OCR-VQA, and VQAv2
- Uses a consistent synthetic voice for the assistant answers, providing a controlled audio environment
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download
- Column-level documentation is absent; field semantics must be inferred after download
Provenance
- Source
- kyutai
- Collection Method
- Text question-answer pairs from three benchmarks were reformatted into conversational dialogue and converted using a text-to-speech pipeline.
- Freshness
- Last updated 2025-03-21 11:47:30