10,738 unique music clips paired with 32,214 images across 13 distinct emotional categories derived from subjective experience studies. Each music entry is linked to three images sharing the same emotional label to support cross-modal sentiment analysis and alignment tasks.
Use Cases
- Train cross-modal retrieval systems to fetch images based on the emotional category of an input music clip
- Develop emotion classification models for audio using the 13-dimensional label set
- Evaluate the emotional consistency of AI-generated imagery against specific musical moods
Strengths
- 10,738 unique music clips and 32,214 associated images
- 13 emotional categories based on the 'What music makes us feel' taxonomy
- Structured as triplets where one music clip corresponds to three images in the same emotional class