Sign in to view source links and access this dataset
Description
25,000 hours of transcribed English speech form the core of this dataset for automatic speech recognition research. The collection includes read and spontaneous speech in both clean and noisy acoustic conditions, organized into subsets of varying size. SpeechBrain authored the dataset, which was last updated on February 11, 2026.
Use Cases
Training large-scale English speech recognition models based on the 25,000-hour 'large' subset.
Benchmarking ASR model robustness based on the described mix of clean and noisy audio conditions.
Developing models for specific speech types based on the described read and spontaneous speech categories.
Creating smaller, curated training sets based on the 2,500-hour 'medium' or 250-hour 'small' subsets.
Strengths
The 'large' subset contains 25,000 hours of transcribed speech, providing substantial scale for training.
The dataset explicitly includes heterogeneous conditions: both read and spontaneous speech, and both clean and noisy audio.
It offers tiered subsets ('large', 'medium', 'small') with different total durations, likely facilitating different research needs.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count and file formats are unknown, which may limit suitability assessment.
The description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
SpeechBrain
Freshness
Last updated 2026-02-11 14:31:57; freshness should be verified.
License restrictions are unknown and must be verified before commercial use.