A dataset likely containing aligned text and audio for Yoruba speech synthesis, sourced from the Bible. It is published on Kaggle, but the author, organization, and creation date are unknown. The specific size, format, and alignment methodology require verification after download.
Use Cases
- Training a speech synthesis model for the Yoruba language (inferred from domain, verify after download)
- Developing forced alignment tools for audio and text corpora (inferred from domain, verify after download)
- Creating pronunciation dictionaries or linguistic resources for Yoruba (inferred from domain, verify after download)
Strengths
- Published on the Kaggle platform, facilitating community access and sharing.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license are unknown, which may limit suitability assessment.