Tobydata TTS Dataset is a Luganda text-to-speech collection created by Bateesa. It contains read speech recordings, primarily on topics related to tailoring, fashion, and vocational training. The data was recorded via a mobile application.
Use Cases
- Train Luganda text-to-speech models based on the described read speech audio.
- Fine-tune speech synthesis models for vocational training topics based on the described content.
- Develop or evaluate automatic speech recognition systems for Luganda based on the audio recordings.
- Create educational or training audio materials for tailoring and fashion in Luganda based on the dataset content.
Strengths
- Contains speech data for Luganda (lg), a lower-resourced language.
- Audio recordings were made via a mobile application, suggesting a potentially accessible collection method.
- Data is focused on specific domains: tailoring, fashion, and vocational training.
Limitations
- Dataset size, row count, and file formats are unknown, which limits suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- TericLab
- Collection Method
- Recorded via mobile application.
- Freshness
- Last updated 2026-02-23 14:56:06; freshness should be verified.