A dataset likely containing aligned text and audio for Yoruba Bible passages, intended for text-to-speech (TTS) applications. It was published on Kaggle, but details on its size, creation date, and author are unknown. The 'ECC' in the title may refer to a specific corpus or project.
Use Cases
- Training a text-to-speech model for Yoruba (inferred from domain, verify after download)
- Evaluating forced alignment algorithms on low-resource languages (inferred from domain, verify after download)
- Creating a Yoruba speech corpus for linguistic research (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing machine learning datasets.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and column definitions are unknown, which limits suitability assessment.
- Data may reflect bias inherent to its specific religious or textual source.