Aligned audio and text data for the Yoruba language, likely derived from the biblical book of 3 John (3JN). The dataset is published on Kaggle, but its specific size, creation date, and author are unknown. Columns suggest it contains paired audio recordings and corresponding text transcripts.
Use Cases
- Training a speech synthesis model for Yoruba (inferred from domain, verify after download)
- Evaluating forced alignment algorithms on a new language (inferred from domain, verify after download)
- Creating a pronunciation dictionary for Yoruba (inferred from domain, verify after download)
Limitations
- Metadata is minimal; actual content requires verification after download
- Row count is unknown, which may limit suitability assessment
- Column-level documentation is absent; field semantics must be inferred after download