Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,575 datasets
arg-spanish-tts is a unified, deduplicated speech corpus for Argentine Spanish (es-AR) containing 10,747 audio rows. The dataset was created by Kukedlc, who merged three public datasets and stripped cross-source duplicates. All audio is resampled to 24 kHz mono, totaling 12.18 hours from 65 unique speakers.
EvA Open Data provides audio clips paired with descriptive captions and instruction-based question-answer data. The audio is sourced from the AudioSet Strong Labels dataset and stored in parquet shards. The dataset was authored by SatsukiVie and last updated on Hugging Face in May 2026.
A dataset titled 'Keyword Spotting Others' was published on the Hugging Face platform by Sarvina. The dataset's specific content and scale are not detailed in the provided metadata. Its last recorded update was on July 6, 2026.
DNS5 Challenge data is a mirrored collection of audio files for speech enhancement tasks. It contains 245 hours of English, 95 hours of French, and 137 hours of German speech sourced from LibriVox, AudioSet, Freesound, OpenSLR26, and OpenSLR28. The dataset was converted to Opus format by user philgzl and last updated in May 2026.
ORNL_CLOUD provides the Site Averaged Flux Data: 1987 (Betts) Data Set from the 1987-1989 FIFE experiment. This dataset contains site-averaged product data collected by multiple principal investigators, structured in 30-minute time intervals for 1987 and covering the entire 1987-1989 period. The data is available in multiple file formats including HTML, PDF, PNG, BIN, ISO, ZIP, and TEXT.
260,162 audio-transcription pairs totaling 730 hours of speech data from 13,290 distinct speakers. This Arabic portion of the YodaLingua collection is designed for training text-to-speech and automatic speech recognition models. The dataset was created by Thomcles and was last updated on May 12, 2026.
An original dataset documents the permitting processes for locally-permitted wind and solar energy projects in Massachusetts. Created by Natalie Baillargeon, it contains data on permitting durations, project outcomes, and capacity. The dataset was last updated on April 8, 2026, and is shared under a CC-BY-4.0 license.
ChildTalk is a large-scale, publicly available multi-dialect Chinese child speech dialogue dataset. It addresses limitations in existing corpora, such as small size and lack of natural conversations, by providing full-length dialogue recordings. The dataset was created by yujie-ovo and was last updated on May 29, 2026.
A dataset for the Mixed Automatic Speech Recognition subtask of the NADI2026 shared task, created by UBC-NLP. The dataset was last updated on June 24, 2026. The specific content and size are not detailed in the provided metadata.
SonoroNova-ES is a large-scale synthetic English-to-Spanish speech-to-speech translation dataset containing 329,764 utterances. It was constructed via cascade pipelines combining text-to-text translation models with neural text-to-speech engines, using source audio derived from the HiFiTTS-2 English audiobook corpus. The dataset features 1,315 unique speakers and provides a total of 961 hours of audio.
A Ukrainian-language speech dataset parsed from the Телебачення Торонто YouTube channel. Each sample consists of a short audio clip paired with its corresponding Ukrainian subtitle text, intended for automatic speech recognition research and education. The dataset was created by yuriilaba and was last updated on Hugging Face in May 2026.
Somali-language audio data published on HuggingFace by jamailyaz. The dataset was last updated on July 8, 2026. Its specific content and scale require verification after download.
Western Australia's Canning Basin provides data on apparent age conflicts in middle Permian stratigraphy. The dataset likely contains U-Pb zircon dates from tuffs and associated palynological zone information, published in a 2017 study by Mory et al. in the Australian Journal of Earth Sciences. It documents a 1.7-million-year discrepancy between CA-IDTIMS dates and established spore-pollen zonation.
Kansas, USA hosts this site-averaged dataset from Portable Automatic Meteorological Stations deployed during the 1987-1989 FIFE experiment. It contains 30-minute interval measurements of atmospheric and surface conditions. The dataset is provided by the National Aeronautics and Space Administration.
Audio segments derived from the VoxCeleb2 dataset, which is a collection of speech from celebrity interviews. The dataset is hosted on Hugging Face by the author AudioJoe and was last updated in July 2026. The specific content and scale of this '3S Chunk' version require verification after download.
9,304 spoken Arabic prompts from 93 users interacting with an ASR and LLM-based assistant. The WASIL dataset, created by QCRI, captures in-the-wild interactions across multiple dialects and countries, including explicit user feedback signals like likes, dislikes, and scalar scores. The dataset was last updated on Hugging Face in May 2026.
Vi Asr Cascaded is a dataset for Vietnamese automatic speech recognition, published on the Hugging Face platform by author pnnbao-ump. The dataset was last updated on June 27, 2026, though specific details on its size, format, and content are not provided in the metadata. Its title suggests it is designed for cascaded ASR model training or evaluation.
Voices in the Wild 2M is an automatic speech recognition dataset designed for robustness training and evaluation. The dataset contains audio files grouped by normalized acoustic subset, with fields for file paths and reference transcriptions. It was created by author zhifeixie and last updated on Hugging Face in May 2026.
1987 data from the FIFE experiment provides site-averaged daily neutron probe soil moisture measurements. The dataset contains product data where samples were averaged first for each site and then for each day. It is managed by ORNL_CLOUD and originates from a field campaign conducted from 1987 to 1989.
Kansas site-averaged gravimetric soil moisture data was collected during the 1987-1989 FIFE field campaign. This dataset contains only the 1988 product, where samples were averaged first by site and then by day. The data is managed by the ORNL_CLOUD organization.