An audio dataset for Indonesian Automatic Speech Recognition (ASR) tasks, published on Kaggle. The dataset likely contains speech recordings and corresponding transcriptions. Specific details on size, collection method, and origin are not provided in the available metadata.
Use Cases
- Train an Indonesian speech-to-text model (inferred from domain, verify after download)
- Benchmark ASR system performance on Indonesian audio (inferred from domain, verify after download)
- Fine-tune a multilingual speech model for Indonesian (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing datasets.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and column definitions are unknown, which may limit suitability assessment.
- Data may reflect geographic or source bias inherent to its collection method.
Provenance
- Geography
- Indonesia (inferred from title)