Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
1,071 hours of Tajik speech data form this corpus for automatic speech recognition. It combines machine-labeled audio from 41 YouTube channels with gold-standard transcriptions from the FLEURS benchmark. The dataset was created by Peacockery and was last updated on June 12, 2026.
License is unknown; users should verify permissions before use. Data is stored in Hive-partitioned Parquet format with audio as FLAC bytes.