A speech dataset likely containing aligned audio recordings and text from the Yoruba Bible. The dataset is published on Kaggle, but its specific size, creation details, and update history are not provided in the available metadata. Its title suggests it is intended for text-to-speech (TTS) model training and alignment tasks.
Use Cases
- Training a text-to-speech model for the Yoruba language (inferred from domain, verify after download)
- Developing forced alignment algorithms for speech and text (inferred from domain, verify after download)
- Creating pronunciation dictionaries or linguistic resources for Yoruba (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing data science resources.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and column details are unknown, which may limit suitability assessment.
- Data may reflect source bias inherent to the specific Bible translation and recording process used.