A dataset for text-to-speech synthesis, likely containing aligned audio recordings and corresponding text transcripts. The data appears to be sourced from the Yoruba Bible, as suggested by the title. It is hosted on Kaggle, but specific details on its size, creation date, and author are not provided.
Use Cases
- Train a text-to-speech model for the Yoruba language (inferred from domain, verify after download)
- Develop speech alignment algorithms using a religious text corpus (inferred from domain, verify after download)
- Benchmark speech synthesis quality for African languages (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing machine learning datasets.
- The title suggests alignment between audio and text, a key feature for TTS training.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and license are unknown, which limits suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Geography
- Likely focused on the Yoruba language, spoken primarily in Nigeria and neighboring West African regions.