Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
Aggregating six indices measuring the economic impact of 109 Music Zones in the United States. The indices assess venue concentration, tourism proximity, business counts, non-chain business presence, total annual economic output, and supported employment.
A dataset titled 'Tts Emotional' published on the Hugging Face platform by SeifElden2342532. The dataset was last updated on March 3, 2026. Its title suggests it likely contains audio data for text-to-speech synthesis with emotional attributes.
A sample of restaurant market data from BeamStation, focusing on technology-ready establishments within Massachusetts, United States. The dataset is a free sample, but the total number of rows, columns, and specific collection date are not provided. The original author and organization are unknown.
A large-scale dataset for deepfake speech detection, created by the CodecFake organization and released in 2025. It includes the CoRS and CoSG subsets, providing audio samples and corresponding protocol and label files for research in synthetic audio generation and detection.
India-focused conversational speech data for evaluating Automatic Speech Recognition systems on Hindi-English code-mixed utterances. The dataset was curated by soketlabs and last updated on the Hugging Face platform in January 2026. It focuses on natural bilingual contexts where Hindi in Devanagari script and English in Latin script co-occur within the same utterance.
A speaker similarity dataset created by thezholdoshbekov, hosted on Hugging Face and last updated in March 2026. The dataset is structured in a tabular format and includes text modality, as indicated by platform tags. It is designed for tasks involving the comparison and identification of speaker voices.
Over 100,000 audio samples for text-to-speech applications, hosted on Hugging Face by datadriven-company. The dataset includes text and corresponding high-fidelity speech audio. It was last updated in March 2026.
Bahraini Speech Dataset is a Bahraini Arabic speech corpus built from publicly available podcast and video content. It contains 90,421 single-speaker utterance clips with aligned transcriptions, created by Hishambarakat and last updated on January 23, 2026.
SPGISpeech is a monolingual English dataset for automatic speech recognition tasks. The dataset is categorized as containing between 1 million and 10 million data instances. It was created by the author 'kensho' and was last updated in January 2026.
An audio dataset titled 'universe-merged-withzero-noASR' is hosted on Kaggle. The dataset's specific content, scale, and creation details are unknown from the provided metadata. Its title suggests it may involve merged audio data, possibly excluding automatic speech recognition (ASR) components.
ASR Full Bundle likely contains audio data for training automatic speech recognition systems. The dataset is hosted on Kaggle, but its specific contents, size, and origin are unknown. Users must download the dataset to verify its actual scope and quality.
ASR-50hour_chunk of lipighor is a dataset for automatic speech recognition (ASR) tasks, published on Kaggle. The title suggests it contains approximately 50 hours of audio data, likely segmented into chunks. The dataset's specific source, collection method, and detailed contents require verification after download.
Encompassing audio recordings of sung poetry from the Pamir Mountains in Tajikistan's Gorno Badakhshan Autonomous Region, collected during fieldwork in 1998. The recordings were made by Jan van Belle and are part of a larger collection spanning multiple years.
A collection of human non-speech vocal sound datasets in WebDataset format, useful for audio classification tasks. The collection includes the 'NonSpeech7k' subset with 7,014 samples across 7 classes like breathing and laughing, sourced from Zenodo. The dataset was authored by 'gijs' and last updated on Hugging Face in January 2026.
Survey responses measuring sleep quality using the Pittsburgh Sleep Quality Index (PSQI). The data is sourced from Kaggle and likely contains self-reported assessments from healthy individuals. Specific details on the number of records, collection period, and original authors are not provided in the metadata.
Muse contains 116,000 synthetic music tracks in Chinese and English, synthesized using SunoV5 and paired with automatically generated lyrics and style descriptions. Created by bolshyC and introduced in early 2026, the collection supports research into reproducible long-form song generation. The data is divided into Chinese (CN) and English (EN) subsets to facilitate multilingual audio modeling.
Serving as from the GAMETES repository, which generates simulated genetic data for studying epistasis. The specific file name suggests it models a 3-way epistatic interaction with 20 attributes and a heritability of 0.2. No row count, column details, or sample data are available.
A release from the GAMETES repository for generating epistasis models. The specific configuration is a 2-way epistasis model with 20 attributes and a heritability of 0.4. Details on row count, columns, and sample data are unavailable.
A GAMETES dataset for epistasis detection, focusing on 2-way interactions with 20 attributes and a heritability of 0.1. The dataset is generated using the EDM-1 model. Specific details on row count, columns, and sample data are unavailable.
A benchmark dataset for Kannada speech recognition tasks, created by thezholdoshbekov. The dataset was last updated in March 2026 and is hosted on the Hugging Face platform with a size category of 1K to 10K entries. It is associated with libraries for tabular and text data processing.