A Yoruba language speech corpus likely intended for text-to-speech applications. The dataset is hosted on Kaggle and appears to be associated with a Jupyter notebook. The number of speakers, recording hours, and specific collection details are unknown.
Use Cases
- Training a text-to-speech model for Yoruba (inferred from domain, verify after download)
- Benchmarking speech synthesis systems on a multi-speaker corpus (inferred from domain, verify after download)
- Conducting linguistic analysis of Yoruba phonetics (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform with integrated data and code sharing.
- Title indicates a multi-speaker corpus, which can improve model robustness.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and audio file specifications are unknown, which may limit suitability assessment.
Provenance
- Source
- Kaggle
- Geography
- Likely Nigeria and other Yoruba-speaking regions.