LibriSpeech-Mixtures: Speech Audio Mixtures for Source Separation
Available on 1 platform
Sign in to view source links and access this dataset
Description
LibriSpeech-Mixtures is an audio dataset hosted on Kaggle, likely derived from the LibriSpeech corpus. The dataset appears to contain mixtures of speech signals, which are commonly used for tasks like source separation. Specific details on the number of files, duration, and creation methodology are not provided in the available metadata.
Use Cases
Training models for blind source separation of overlapping speech (inferred from domain, verify after download)
Benchmarking speech enhancement algorithms on mixed audio signals (inferred from domain, verify after download)
Developing speaker diarization systems in noisy or multi-speaker environments (inferred from domain, verify after download)
Strengths
Published on Kaggle, a platform with established data hosting and versioning.
Platform tags indicate a focus on speech mixtures and audio processing, suggesting domain relevance.
Limitations
Metadata is minimal; actual content, file formats, and scale require verification after download.
Column-level documentation is absent; field semantics must be inferred after download.
License, author, and last update information are unknown, affecting reproducibility and usage clarity.
Provenance
Source
Likely derived from the LibriSpeech corpus, but the specific processing and mixing methodology are unknown.
License terms are unknown; users should verify permissible uses before downloading.