Sign in to view source links and access this dataset
Description
16,017 audio samples filtered from a larger 616-hour speech dataset to contain only ElevenLabs Scribe v1 audio events. The dataset, created by TTS-AGI, focuses on vocal bursts and background sounds with unified annotation formatting. It was last updated on March 28, 2026.
Use Cases
Training TTS models to generate expressive speech based on the included vocal burst annotations.
Augmenting speech datasets with non-verbal sounds based on the described background audio events.
Studying paralinguistic features in synthetic speech based on the curated event annotations.
Strengths
Contains 16,017 curated samples specifically for audio events.
Annotations use a consistent format with square brackets for vocal bursts.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for the original source is known, but the final dataset's exact size beyond the filtered count is unknown.
Provenance
Source
Filtered subset of MrDragonFox/EN_Emilia_Yodas_616h.
Collection Method
Filtered to only include rows where events_scribe is non-empty.
Freshness
Last updated 2026-03-28 12:17:09; freshness should be verified.
License is unknown; terms of use must be verified before application.