Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,571 datasets
Experimental data from the Pitch Imagery Arrow Task investigating the effects of musical training, vividness, and mental control. The dataset includes R code for reproducing figures from the associated research article and three data files. The data and code were contributed by author Rebecca Gelding.
2,702 prompt-target audio pairs for evaluating zero-shot text-to-speech systems across three hierarchical dimensions: basic generalization, paralinguistic control, and hard robustness scenarios. The dataset, created by dinosaaaur, contains approximately 1GB of audio in Chinese and English and was last updated on June 23, 2026. It is structured into three subsets targeting different evaluation goals.