OpenSLR LibriSpeech ASR 12000-15999 is a speech recognition dataset published on the Hugging Face platform by user Kimang18. The dataset, last updated on 2026-07-16, likely contains audio recordings and corresponding text transcripts, as suggested by its title and platform tags. The specific number of samples, file formats, and license details are not provided in the available metadata.
Use Cases
- Train an acoustic model for English speech recognition (inferred from domain, verify after download)
- Benchmark speech-to-text transcription accuracy (inferred from domain, verify after download)
- Fine-tune a pre-trained ASR model on read speech (inferred from domain, verify after download)
Strengths
- Published on the Hugging Face platform, facilitating access and integration with ML tools.
- Based on the established LibriSpeech corpus, a common benchmark in speech recognition.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count, file formats, and license are unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Source
- huggingface
- Freshness
- Last updated 2026-07-16 11:18:48; freshness should be verified.