Yoruba-language audio and text data from the Gospel of John, likely aligned for text-to-speech model training. The dataset is hosted on Kaggle, but its size, creator, and specific structure are not detailed. Columns and sample data are unknown, requiring download for verification.
Use Cases
- Training a text-to-speech model for Yoruba (inferred from domain, verify after download)
- Creating a phoneme-aligned pronunciation dictionary for Yoruba (inferred from domain, verify after download)
- Benchmarking speech synthesis quality on religious domain text (inferred from domain, verify after download)
Limitations
- Metadata is minimal; actual content requires verification after download
- Column-level documentation is absent; field semantics must be inferred after download
- Row count is unknown, which may limit suitability assessment
Provenance
- Geography
- Yoruba language region (inferred from title)