Yoruba-language audio data, likely aligned with biblical text for speech synthesis. The dataset is hosted on Kaggle, but its specific size, creation date, and author are unknown. Columns suggest it contains audio files and corresponding text transcripts.
Use Cases
- Training a Yoruba speech synthesis model (inferred from domain, verify after download)
- Evaluating forced alignment algorithms on religious text (inferred from domain, verify after download)
- Fine-tuning a multilingual TTS system (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing machine learning datasets.
- The title indicates the data is aligned, suggesting a structured pairing of audio and text.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- Kaggle
- Geography
- Likely Nigeria or the Yoruba-speaking region (inferred from language).