Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A curated Urdu speech dataset for Text-to-Speech and Automatic Speech Recognition research. Audio segments are extracted from publicly available YouTube speech content, processed through a multi-stage quality pipeline, and annotated with Urdu transcriptions. This mini release by salisai, last updated in June 2026, is intended to validate the preprocessing pipeline and establish a quality baseline for future large-scale versions.
License information is unknown; terms of use must be verified on the dataset page.