Audio Emotion Detection Dataset contains speech clips in English and Hindi annotated with emotion labels and ASR transcripts. Audio is sourced from public YouTube videos and trimmed to approximately 60 seconds per clip, with noise reduction applied. The dataset was created by RapidOrc121 and last updated on 2026-06-07.
Use Cases
- Train speech emotion recognition models based on the five annotated emotion classes.
- Benchmark automatic speech recognition systems using the provided transcripts.
- Study cross-lingual emotion expression in speech using the English and Hindi clips.
- Develop audio preprocessing pipelines using the described noise reduction and trimming methods.
Strengths
- Clips are processed to a consistent length of approximately 60 seconds.
- Data includes five distinct emotion classes with descriptions.
- Audio preprocessing includes noise reduction and voice activity detection.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Data may reflect bias inherent to the source YouTube videos.
Provenance
- Source
- RapidOrc121 on Hugging Face
- Collection Method
- Sourced from public YouTube videos, trimmed, and processed with noise reduction.
- Freshness
- Last updated 2026-06-07 16:25:54; freshness should be verified.