Sign in to view source links and access this dataset
Description
A high-quality Russian-language audio dataset for Text-to-Speech and Automatic Speech Recognition tasks. It contains 17,670 audio samples totaling 39 hours, 35 minutes, and 29 seconds of speech, processed using a modern pipeline by TeraTTS. The dataset was last updated on February 27, 2026.
Use Cases
Training Text-to-Speech models based on Russian audio samples.
Fine-tuning Automatic Speech Recognition models using Russian speech data.
Benchmarking speech synthesis quality on a dataset with detailed duration statistics.
Analyzing speech characteristics based on sample length distribution.
Strengths
Contains 17,670 individual audio samples.
Provides substantial total audio duration of 39 hours, 35 minutes, and 29 seconds.
Offers detailed sample length statistics, including average (8.07 seconds), median (8.01 seconds), minimum (0.25 seconds), and maximum (24.57 seconds) durations.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Data may reflect source bias inherent to YouTube content.
Provenance
Source
YouTube
Collection Method
Collected and processed using a modern pipeline.
Freshness
Last updated 2026-02-27 15:42:32; freshness should be verified.
Geography
Russian-language content.
License is unknown; terms of use must be verified.