Sign in to view source links and access this dataset
Description
17,476 preprocessed Urdu speech samples from Mozilla Common Voice, split into training, validation, and test sets. The dataset is processed for Whisper models, with audio resampled to 16kHz. It was uploaded by khawajaaliarshad and last updated on 2025-12-27.
Use Cases
Fine-tuning Whisper-based ASR models based on pre-resampled 16kHz audio.
Hyperparameter tuning for speech models based on the provided validation split of 5,046 samples.
Benchmarking ASR model performance based on the dedicated test split of 5,091 samples.
Training speech recognition systems for Urdu based on the training split of 7,339 samples.
Strengths
Dataset is preprocessed and ready for Whisper models, with audio resampled to 16kHz.
Provides standard splits for machine learning: 7,339 training, 5,046 validation, and 5,091 test samples.
Contains a total of 17,476 speech samples for the Urdu language.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
Mozilla Common Voice
Collection Method
Likely contains crowd-sourced speech recordings, processed for ASR use.
Freshness
Last updated 2025-12-27 05:10:18; freshness should be verified.
License information is unknown and should be verified before use.