Yoruba-language audio and text data for the First Book of Samuel (1SA), aligned for text-to-speech (TTS) applications. The dataset is published on Kaggle, but its creator, size, and specific alignment methodology are unknown. Its content likely consists of Yoruba Bible verses paired with corresponding speech recordings.
Use Cases
- Train a text-to-speech model for Yoruba (inferred from domain, verify after download)
- Create a Yoruba speech corpus for linguistic analysis (inferred from domain, verify after download)
- Benchmark forced alignment algorithms on African language data (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform with established data hosting and versioning.
- Focuses on Yoruba, a major West African language, addressing a potential gap in speech resources.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license are unknown, which limits suitability assessment.
Provenance
- Source
- Kaggle
- Geography
- Yoruba language region (inferred from content)