TUT Sound Events 2018: Simulated Ambisonic Recordings with Overlapping Sources
by Sharath Adavanne / Tampere University
Available on 1 platform
Sign in to view source links and access this dataset
Description
240 training and 60 test recordings, each about 30 seconds long, simulate first-order Ambisonic audio with up to three overlapping sound events. The dataset, created by Sharath Adavanne at Tampere University, uses 11 sound event classes from the DCASE 2016 task 2 dataset, spatially placed in a simulated 10x8x4 meter room. Each recording includes metadata with event timings and spatial coordinates in azimuth, elevation, and distance.
Use Cases
Training sound event localization models based on simulated first-order Ambisonic recordings with spatial metadata.
Developing algorithms for detecting overlapping sound events based on datasets with up to three temporally overlapping sources.
Benchmarking acoustic scene analysis systems using synthetic impulse responses and reverberant conditions described in the dataset.
Researching spatial audio perception and source separation using data with annotated azimuth, elevation, and distance coordinates.
Strengths
Provides three distinct sub-datasets for training models with one, two, or three temporally overlapping sound events.
Includes 240 training and 60 test recordings per cross-validation split, each about 30 seconds long.
Metadata for each recording specifies sound event class, onset/offset times, and precise spatial coordinates (azimuth, elevation, distance).
Sound events are placed using a defined spatial grid with 10-degree azimuth resolution and elevation between -60 and 60 degrees.
Limitations
Row count and total file size are unknown, which may limit suitability assessment for large-scale projects.
Column-level documentation is absent; field semantics must be inferred after download.
Data is entirely synthetic, generated using the image source method, which may not fully capture real-world acoustic variability.
Provenance
Source
Tampere University
Collection Method
Synthetic generation using the image source method, with isolated sound events sourced from the DCASE 2016 task 2 dataset.
License details are contained in a separate LICENSE file and must be reviewed before use.