Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
Speech recognition data published on Kaggle. The dataset's specific content, scale, and origin are not detailed in the available metadata. Further inspection after download is required to confirm the actual audio files, transcripts, and recording conditions.
Sam Wake Word is a dataset uploaded to Hugging Face by author sh1vam10. The dataset's platform tags indicate it contains audio and text modalities, likely for wake word or keyword spotting tasks. It was last updated on March 20, 2026.
TAPS: Throat and Acoustic Paired Speech Dataset is a standardized corpus for deep learning-based speech enhancement, specifically targeting throat microphone recordings. The dataset provides paired recordings from 60 native Korean speakers, designed to address the high-frequency attenuation in throat mics caused by the low-pass filtering effect of skin and tissue.
Kaggle hosts a dataset titled 'music_file'. The dataset likely contains audio files related to music. Metadata is minimal; the specific content, scale, and origin require verification after download.
Audio recordings of general utterances feature Hindi speakers from India. The dataset's size, collection date, and creator are not specified in the provided metadata. It is hosted on the Kaggle platform.
A collection of French-language audio recordings from a medical call center. The dataset is hosted on Kaggle and is intended for speech processing tasks in the healthcare sector. Specific details on size, creation date, and authorship are not provided.
Sarvam AI developed this synthetic benchmark in 2026 to evaluate context-aware Automatic Speech Recognition (ASR) within voice bot environments. The collection includes between 1,000 and 10,000 records covering the top 10 Indian languages, focusing on how conversation history and agent prompts influence transcription accuracy.
An audio dataset likely contains general utterances spoken by Spanish speakers from Spain. The dataset's size, specific content, and creation details are unknown. It is hosted on Kaggle.
A high-quality, single-speaker Persian (Farsi) narration dataset intended for training text-to-speech models. The dataset was created by author pymmdrza and was last updated on January 21, 2026. The description emphasizes professional narration quality for TTS applications.
Updated Hate-Speech Dataset is a text corpus likely containing social media posts or comments annotated for offensive language. The dataset is hosted on Kaggle, but its specific size, origin, and update details are not provided in the metadata. Columns and sample data are unknown, requiring verification after download to confirm content and structure.
An audio dataset named TrainXttsV2_Audiobook, likely containing speech recordings for text-to-speech model development. The dataset is hosted on Kaggle, but its specific size, creator, and update date are unknown. Columns and sample data are unavailable, so the exact content requires verification after download.
Tts Polish Nemo is a dataset for text-to-speech synthesis, published on HuggingFace by datadriven-company. The dataset was last updated on March 13, 2026. Its specific content and scale require verification after download.
A dataset from the OpenML platform with an identifier suggesting it relates to genetic heterogeneity modeling. No concrete details on size, content, or structure are available from the provided input.
A dataset from the GAMETES repository for generating epistasis models. The specific attributes, sample size, and data structure are unknown.
BIGOS (Benchmark Intended Grouping of Open Speech) is a collection of openly available Polish speech corpora. Its goal is to simplify access to these resources and enable systematic benchmarking of open and commercial Polish automatic speech recognition (ASR) systems. The dataset was created by amu-cai and was last updated on 2026-02-18.
A dataset titled 'fasrtyuj' is available on the Kaggle platform. The dataset's content, structure, and origin are not described in the provided metadata. Further details about its creation, size, and specific contents require verification after download.
MusicOne is a dataset hosted on Kaggle. Its title suggests a focus on music-related information. The dataset's specific content, scale, and origin require verification after download due to minimal provided metadata.
MusicSecond is a dataset hosted on Kaggle. Its title suggests it contains audio data related to music. The dataset's specific content, size, and origin are not detailed in the available metadata.
A dataset titled 'MusicThree' published on Kaggle. The dataset's content likely relates to music, but specific details such as size, format, and creation date are unavailable. Metadata is minimal; actual content requires verification after download.
Kazakh language speech data comprising approximately 726 hours of audio in FLAC format at a 16kHz sampling rate. The dataset is designed to support Automatic Speech Recognition and Text-to-Speech system development. It is an open-source corpus created by Flamme-VRM and was last updated on Hugging Face in January 2026.