Yoruba BibleTTS Aligned - GEN is a dataset published on Kaggle. It likely contains audio recordings and corresponding text transcripts aligned for speech synthesis tasks. The dataset's specific size, origin, and creation date are not provided in the available metadata.
Use Cases
- Training a text-to-speech model for the Yoruba language (inferred from domain, verify after download)
- Developing forced alignment tools for audio and text corpora (inferred from domain, verify after download)
- Creating pronunciation dictionaries or linguistic resources for Yoruba (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform with established data sharing and versioning.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and file size are unknown, which may limit suitability assessment.