A text-to-speech dataset for the Yoruba language, published on Kaggle. It is part of the MMS (Massively Multilingual Speech) project by Meta (formerly Facebook). The dataset likely contains audio samples paired with corresponding text transcripts.
Use Cases
- Training a text-to-speech model for the Yoruba language (inferred from domain, verify after download)
- Benchmarking multilingual speech synthesis systems (inferred from domain, verify after download)
- Fine-tuning a pre-trained speech model for a specific Yoruba dialect or accent (inferred from domain, verify after download)
Strengths
- Published on Kaggle.
- Part of the Meta MMS project, suggesting a research-grade origin.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and file size are unknown, which may limit suitability assessment.
Provenance
- Source
- Meta (Facebook) MMS project