TTS-Hungarian is a large-scale speech dataset containing 253,116 audio samples totaling 702 hours, derived from the Magyar Elektronikus Könyvtár (MEK) collection of Hungarian audiobooks. It features recordings from 100 unique speakers, with an average sample duration of 10.0 seconds and an average DNSMOS quality score of 3.68. The dataset was created by the author 'datadriven-company' and was last updated on the Hugging Face platform in February 2026.
Use Cases
- Train text-to-speech models based on high-quality Hungarian speech samples.
- Develop automatic speech recognition systems based on Hungarian audiobook recordings.
- Benchmark speech synthesis quality using the provided DNSMOS scores.
- Study speaker characteristics and variation based on the 100 unique speakers.
Strengths
- Large scale with 253,116 samples and 702 hours of total audio.
- High-quality source material derived from professional Hungarian audiobooks.
- Includes 100 unique speakers, providing diversity for model training.
- Quantified audio quality with an average DNSMOS score of 3.68.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- Magyar Elektronikus Könyvtár (MEK) — Hungarian audiobooks.
- Collection Method
- Derived from existing audiobook collections.
- Freshness
- Last updated 2026-02-12 10:10:41; freshness should be verified.
- Geography
- Hungary (implied by language and source)