Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
2502_dataset_TTS is a Kaggle-hosted collection likely containing audio data for text-to-speech applications. The dataset's specific content, size, and origin are unconfirmed due to minimal metadata. Its title suggests it may include speech samples or synthesis parameters for machine learning model training.
A dataset likely focused on the recognition of musical notes from audio signals. The title suggests it includes a cross-validation scheme, which may indicate a structured setup for model evaluation. It is published on Kaggle, but details on its size, origin, and specific content are unavailable.
A dataset titled 'hinglish-tts-src' is hosted on Kaggle. The title suggests it contains source materials for text-to-speech synthesis in Hinglish, a code-mixed language of Hindi and English. The dataset's specific content, size, and creation details are unknown from the provided metadata.
Kaggle hosts the RTTS_C0C0 dataset. The title suggests it may relate to a specific project or codename. Its content and structure require verification after download.
A 235-mile polyline feature depicting the New England National Scenic Trail from the Long Island Sound in Guilford, Connecticut, to the Massachusetts/New Hampshire border. The dataset was created by combining work from the Connecticut Forest & Park Association and the Appalachian Mountain Club. It was last updated on March 4, 2026.
Kikuyu language audio data, preprocessed for automatic speech recognition tasks. The dataset was published on huggingface by the author InterstellarCG and was last updated on March 16, 2026. The specific content, scale, and preprocessing methods require verification after download.
Dysarthric speech data published on Kaggle. The dataset likely contains audio recordings of speech affected by motor speech disorders. Specific details on size, collection method, and origin are not provided in the available metadata.
Cleaned Asr Transcripts is a text dataset published on Hugging Face by author bingbangboom. The dataset likely contains processed transcripts generated by an Automatic Speech Recognition (ASR) system. It was last updated on March 24, 2026.
VoxCeleb2 is an audio dataset published on Hugging Face by the author 'humanify'. The dataset was last updated on March 25, 2026. Its specific content, size, and license details are not provided in the available metadata.
Audio data related to music, likely intended for training machine learning models. The dataset is hosted on Kaggle, but its specific contents, size, and creation details are not provided in the metadata. Users must download the dataset to verify its exact composition and quality.
Large-scale CC0 Pashto speech dataset for Automatic Speech Recognition (ASR). The dataset is part of the Common Voice project, version 25.0, and is hosted on Kaggle. Its specific collection method, size, and contributor details are not provided in the available metadata.
A collection of pairwise human preferences for music generated by text-to-music systems. The dataset is hosted on Huggingface Datasets by the author 'i-need-sleep' and was last updated on 2026-01-27. It is intended for research into evaluating and improving generative music models.
Synthetic audio data generated based on the VS13 framework, likely containing simulated vehicle sounds. The dataset is hosted on Kaggle, but details on its size, creation method, and specific contents are not provided. Metadata is minimal; actual content requires verification after download.
An AI-generated voice dataset for the Nepali language, published on Kaggle. The dataset is likely designed for text-to-speech (TTS) synthesis, modeled after the LJ Speech dataset structure. Its specific size, creation date, and author details are not provided in the available metadata.
A multilingual automatic speech recognition dataset covering 30 Indic dialects and languages. It contains over 2.8 million audio samples with corresponding transcriptions. The dataset was created by author grushaaaaa and last updated on Hugging Face in February 2026.
UniDataPro's collection features 338 hours of Russian telephone dialogues recorded from 460 native speakers across diverse topics. Updated in January 2026, the data is specifically formatted for automatic speech recognition (ASR) research and model training. It maintains a verified 98% Word Accuracy Rate for its transcriptions.
Pulse 2026 is a high-fidelity synthetic music dataset with engineered streaming metrics. The dataset appears to focus on music evolution and viral analytics. Its specific source, size, and creation date are unknown.
Music-Gen-Task3-Split is a dataset hosted on Kaggle, likely related to a music generation challenge. The title suggests it contains audio data split for a specific machine learning task, though the exact content and structure are unspecified. No information is available regarding its author, size, or creation date.
Hate speech detection data spanning two major languages, English and Spanish. The dataset is hosted on Kaggle, but its specific collection method, size, and annotation details are not provided in the available metadata. Researchers must download the dataset to inspect its volume, annotation schema, and source characteristics.
This Russian speech corpus contains audio recordings across diverse genres including podcasts, public speeches, YouTube content, audiobooks, and phone calls. The dataset was processed using the BALALAIKA pipeline by the MTUCI lab260 team to provide high-quality annotations for generative speech tasks.