2,702 prompt-target audio pairs for evaluating zero-shot text-to-speech systems across three hierarchical dimensions: basic generalization, paralinguistic control, and hard robustness scenarios. The dataset, created by dinosaaaur, contains approximately 1GB of audio in Chinese and English and was last updated on June 23, 2026. It is structured into three subsets targeting different evaluation goals.
Use Cases
- Benchmarking speaker similarity and text accuracy based on the basic-v1 subset's core evaluation targets.
- Evaluating paralinguistic control for 18 types of non-linguistic vocalizations based on the paralinguistic-v1 subset.
- Testing model robustness on long sentences, poetry, proper nouns, and code-switching based on the hard-v1 subset's high-difficulty scenarios.
Strengths
- Contains 2,702 audio pairs, providing a substantial corpus for evaluation.
- Structured into three hierarchical subsets (basic-v1, paralinguistic-v1, hard-v1) with 1,102, 1,106, and 494 samples respectively, enabling targeted testing.
- Covers both Chinese and English languages, with specific sample counts per subset and language.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Last updated 2026-06-23 08:17:00; freshness should be verified.
Provenance
- Source
- huggingface user dinosaaaur
- Freshness
- Last updated 2026-06-23 08:17:00.