A dataset for text-to-speech synthesis in the Yoruba language, likely containing aligned audio and text. It is hosted on Kaggle, but the specific creation date and author are unknown. The dataset's size, exact format, and number of samples are unspecified.
Use Cases
- Training a text-to-speech model for Yoruba (inferred from domain, verify after download)
- Developing pronunciation dictionaries or grapheme-to-phoneme models (inferred from domain, verify after download)
- Benchmarking speech synthesis quality for African languages (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing machine learning datasets.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and column definitions are unknown, which limits suitability assessment.
- Data may reflect source or collection bias inherent to the original compilation.