LibriSpeech8K_100_360: Speech Audio Corpus for ASR
Available on 1 platform
Sign in to view source links and access this dataset
Description
LibriSpeech8K_100_360 is a speech audio dataset published on Kaggle. The title suggests it is derived from the LibriSpeech corpus, likely containing 8,000 audio samples. The specific content, such as speaker count, recording length, and transcription details, requires verification after download.
Use Cases
Training an acoustic model for speech-to-text conversion (inferred from domain, verify after download)
Benchmarking ASR system performance on clean read speech (inferred from domain, verify after download)
Fine-tuning pre-trained models for speaker or accent recognition (inferred from domain, verify after download)
Strengths
Published on Kaggle, a major platform for data science resources.
Limitations
Metadata is minimal; actual content requires verification after download.
Column-level documentation is absent; field semantics must be inferred after download.
Row count, file formats, and license are unknown, which may limit suitability assessment.
Provenance
Source
Likely derived from the LibriSpeech corpus.
Collection Method
Method of gathering is unknown.
Time Range
Temporal coverage is unknown.
Freshness
Last updated date is unknown; freshness unverified.
Geography
Spatial coverage is unknown.
License is unknown; users must verify permissions before commercial use.