Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
6,974 audio samples of 15 to 30 seconds duration, each paired with five captions of 8 to 20 words, resulting in 34,870 total captions. The dataset was created by Konstantinos Drossos of Tampere University and is described in a 2020 ICASSP paper. It is structured into development, validation, and evaluation splits.
Use requires citing the associated paper. The dataset files are split into development, validation, and evaluation archives and CSV files.