Adaption Low Resource Audio is a subset of the PolyglotAudio dataset, remastered with Adaption's Adaptive Data platform. It contains 3,704 rows of paired audio clips and text, spanning 10 languages typically underrepresented in open corpora. The dataset was created by Reubencf and last updated on April 24, 2026.
Use Cases
- Fine-tuning speech recognition models based on paired audio clips and text.
- Evaluating text-to-speech systems on underrepresented languages based on the audio-text pairs.
- Training multilingual speech models based on data spanning 10 languages.
- Benchmarking model performance on low-resource-language tasks based on the enhanced prompt/completion columns.
Strengths
- Contains 3,704 rows of paired audio and text data.
- Focuses on 10 languages that are typically underrepresented in open ASR/TTS corpora.
- Includes enhanced_prompt and enhanced_completion columns, likely prepared for direct use in model training.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is known, but other size metrics (file formats, total size) are unknown.
- Freshness should be verified; the last update date is April 24, 2026.
Provenance
- Source
- Reubencf/PolyglotAudio subset remastered with Adaption's Adaptive Data platform.
- Collection Method
- Derived from Tatoeba audio clips, with added enhanced columns.
- Freshness
- Last updated 2026-04-24 10:33:15.