Sign in to view source links and access this dataset
Description
WazobiaVoice TTS (wazobia-tts-cc) is a transcribed speech dataset containing 5,866 audio clips totaling approximately 18.0 hours. It covers five Nigerian language varieties: Yoruba, Hausa, Igbo, Nigerian-accented English, and Nigerian Pidgin. The dataset was created by Axiveri and is licensed under CC-BY-4.0.
Use Cases
Train text-to-speech models based on the transcribed short audio clips in five language varieties.
Develop speech recognition systems for Nigerian languages based on the transcribed audio data.
Fine-tune language models for Nigerian Pidgin or accented English using the speech transcriptions.
Conduct linguistic analysis of Nigerian language varieties based on the parallel speech and text data.
Strengths
Contains 5,866 quality-filtered audio clips, providing a substantial corpus for model training.
Covers five distinct language varieties, including three major Nigerian languages and two English variants.
Provides language distribution statistics, with Hausa comprising 40.2% (2,360 clips) of the dataset.
Audio is transcribed and filtered for language accuracy, suggesting curated linguistic content.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
The dataset consists of single-speaker clips, which may limit speaker diversity for some applications.
Row count is unknown, which may limit suitability assessment for large-scale training.
Provenance
Source
Axiveri on Hugging Face.
Collection Method
Likely contains single-speaker short clips that were transcribed and filtered for quality and language accuracy.
Freshness
Last updated 2026-07-05 11:44:23; freshness should be verified.