Yoruba BibleTTS Aligned - DAN is a dataset for text-to-speech research, likely containing audio recordings aligned with text transcripts. The dataset's specific size, structure, and creation details are unknown. It is hosted on Kaggle, a platform for sharing machine learning datasets.
Use Cases
- Train a text-to-speech model for the Yoruba language (inferred from domain, verify after download)
- Benchmark forced alignment algorithms on a new language corpus (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a major platform for sharing machine learning data.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, column definitions, and file formats are unknown, which may limit suitability assessment.
- Data may reflect source bias inherent to the specific Bible translation and recording methodology used.