Kaggle hosts this audio dataset derived from the LibriSpeech corpus. The title suggests it contains speech recordings with added noise, intended for training or evaluating MetricGAN+ and Automatic Speech Recognition systems. The dataset's author, organization, and specific details like size and license are unknown.
Use Cases
- Training a speech enhancement model to denoise audio (inferred from domain, verify after download)
- Benchmarking ASR system performance on noisy speech (inferred from domain, verify after download)
- Synthesizing noisy audio samples for data augmentation (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a major data science platform.
- Based on the well-known LibriSpeech corpus, a standard in speech research.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count and file size are unknown, which may limit suitability assessment.
Provenance
- Source
- LibriSpeech corpus