Yoruba language audio and text data, likely aligned for speech synthesis tasks. The dataset is published on Kaggle, but its specific size, creation date, and author are unknown. Its content appears to be derived from Biblical text, suggesting a focus on formal or religious speech.
Use Cases
- Training a speech synthesis model for Yoruba (inferred from domain, verify after download)
- Aligning text transcripts with audio waveforms for prosody analysis (inferred from domain, verify after download)
- Creating a pronunciation dictionary for Yoruba speech processing (inferred from domain, verify after download)
Strengths
- Published on the Kaggle platform, facilitating community access and sharing.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license are unknown, which may limit suitability assessment.
Provenance
- Source
- Kaggle
- Geography
- Likely contains Yoruba language content, which is primarily spoken in Nigeria and neighboring West African regions.