MAESTRO Synthetic: 20 Synthetic Soundscapes for Crowdsourced Label Estimation
by Irene Martín-Morató / Tampere University
Available on 1 platform
Sign in to view source links and access this dataset
Description
20 synthetic audio files, each 3 minutes long, were created using the Scaper tool to study the estimation of strong labels from crowdsourced annotations. The dataset includes reference annotations, crowdsourced estimated strong labels, and weak labels for each 10-second segment, generated via Amazon Mechanical Turk. It was created by researchers at Tampere University for a 2021 WASPAA paper on crowdsourcing strong labels for sound event detection.
Use Cases
Training sound event detection models based on the provided synthetic soundscapes and reference annotations.
Developing algorithms for estimating strong labels from crowdsourced data based on the included annotation outcomes.
Studying the relationship between weak and strong labels using the per-segment audio tags.
Benchmarking label aggregation methods for audio based on the Mechanical Turk annotation results.
Strengths
Contains 20 distinct audio files, each with a controlled 3-minute duration.
Provides three layers of annotation: ground truth, estimated strong labels, and per-segment weak labels.
Synthetic nature allows for controlled study of label estimation without real-world recording variability.
Limitations
Dataset size is limited to 20 files; scale may be insufficient for training large models.
Column-level documentation is absent; field semantics must be inferred after download.
Data is synthetic, which may limit generalizability to real-world audio recordings.
Provenance
Source
Tampere University
Collection Method
Audio files were synthetically generated using Scaper, with annotations crowdsourced via Amazon Mechanical Turk.
Audio excerpts are derived from the Urban Sound 8k dataset via freesound.org; users must check the FREESOUNDCREDITS.txt file for attribution.