Sign in to view source links and access this dataset
Description
Chalermdej's Yodas2 Sidon Th Tts is a filtered, quality-verified Thai text-to-speech dataset derived from sarulab-speech/yodas2_sidon. It contains 141,927 audio samples totaling 156.0 hours from 4,199 speakers, with transcriptions verified by multiple ASR models and Gemini. The dataset was last updated on 2026-06-06.
Use Cases
Train text-to-speech models based on the 141,927 quality-screened Thai audio samples.
Benchmark ASR model performance on Thai speech based on the multi-model verified transcriptions.
Develop voice cloning or multi-speaker TTS systems based on the 4,199 distinct speaker identities.
Study Thai speech prosody and phonetics based on the text-normalized and quality-verified audio corpus.
Strengths
Large scale with 141,927 audio samples and 156.0 hours of speech.
Quality verification includes ASR and Gemini-based transcription checks and DNSMOS audio quality screening.
Text normalization to Thai ensures linguistic consistency for model training.
Multi-speaker diversity with contributions from 4,199 speakers.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
License information is unknown, which may restrict usage.
Data may reflect source bias inherent to the original YODAS2 sidon collection.
Provenance
Source
Derived from sarulab-speech/yodas2_sidon on Hugging Face.
Collection Method
Filtered and quality-verified by the author Chalermdej, with transcriptions verified by multiple ASR models and Gemini.
Freshness
Last updated 2026-06-06 03:57:12; freshness should be verified.
License restrictions are unknown and must be verified before use.