Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,575 datasets
Oceanographic measurements collected in Massachusetts Bay and the surrounding area. The dataset covers a multi-year period from 2002 to 2005 and is provided by the National Aeronautics and Space Administration. Data includes parameters related to ocean chemistry, optics, temperature, and salinity.
NASA collected in-situ oceanographic data along the coastal regions of New Hampshire and Massachusetts during 2009. The dataset includes measurements related to ocean chemistry, optics, temperature, and salinity. It is available in BIN and ISO file formats.
SpeakerCard-1M is a speaker-centric corpus built on the VoxCeleb1 and VoxCeleb2 datasets. It was created by JYP2024 using a tool-first, LLM-last pipeline where ten acoustic probes extract evidence for a structured schema. The dataset was last updated on June 3, 2026.
Tajik ASR Corpus v0 is a deduplicated collection for automatic speech recognition assembled from multiple sources. The dataset, created by Peacockery, includes data from FLEURS-derived speech, Mozilla Common Voice 25 Tajik, and augmented data from Muhtasham Tajik ASR. Each data split is provided in TSV format with an audio directory, and a SQLite version includes additional normalized fields.
OpenSLR LibriSpeech ASR 12000-15999 is a speech recognition dataset published on the Hugging Face platform by user Kimang18. The dataset, last updated on 2026-07-16, likely contains audio recordings and corresponding text transcripts, as suggested by its title and platform tags. The specific number of samples, file formats, and license details are not provided in the available metadata.
RGAD Cross-Lingual TTS 10h is a 10-hour speech dataset for prompt-conditioned text-to-speech fine-tuning. It was created by author 'isabeth' and last updated on Hugging Face in May 2026. The dataset contains paired audio prompts and targets across languages, specifically for generating Chinese speech from prompts in other languages.
Ghana is the primary source for KasaSpeech, a large-scale speech dataset featuring natural switching between English and Twi. It contains 49,878 transcribed audio samples, split into training, validation, and test sets. The dataset was created by Kennethdot and last updated on Hugging Face in May 2026.
A domain-specific Pashto automatic speech recognition dataset covering agriculture, general topics, food services, health, and services. The dataset is structured by domain with audio files and corresponding transcript CSV files, created by Sabtain-Dev and last updated on June 5, 2026.
Mory et al. (2017) published zircon U–Pb ages from middle Permian tuffs in Western Australia's Canning Basin. The data reveals an apparent conflict between CA-IDTIMS ages and established spore-pollen zonation, with a specific age of 267.04 ± 0.14 Ma reported from the Pittston SD-1 drillhole. The dataset is hosted by the Australian Ocean Data Network and was last updated in April 2026.
MERIT is a dataset of audio triplets designed for training a framework that learns three independent music similarity spaces: melody, rhythm, and timbre. It was created by the AMAAI-Lab and is hosted on Hugging Face. The dataset page was last updated on 2026-05-26.
Between September 20 and December 27, 2001, a Rapid Single-Particle Mass Spectrometer (RSMS) captured real-time composition data for individual aerosol particles in Pittsburgh. Each record includes aerodynamic particle size, positive and negative mass spectra, and precise measurement time, enabling analysis of particle-to-particle variation. The data covers nine logarithmically spaced size classes from about 40 to 1300 nanometers.
An anonymised dataset from a study investigating music performance anxiety and flow under performance simulation conditions. The dataset is 49.9 KB in size and was last updated on 2026-05-21. It was published by a research team under a CC-BY-4.0 license on figshare.
Primary survey results from a post-event questionnaire conducted in the coastal region of Toyama Prefecture, Japan, following the 2024 Noto Peninsula Earthquake tsunami. The dataset was created by Shuichi Kure and is associated with a 2025 research paper in Coastal Engineering Journal. It consists of 1.4 MB of data available in PDF, TXT, and XLSX formats.
Establishments of the Conservatory of Music and Dramatic Art of Quebec provides a list and geolocation of its establishments. The dataset is published by the Government and Municipalities of Québec under a CC-BY-4.0 license and was last updated on 2026-04-22.
A codebook for analyzing storytelling in music content on TikTok. The dataset is published on the Papers with Code platform under an Open Access license. The author is listed as 'a v', but other details like size and update date are unknown.
100 high-fidelity, simulated clinical teleconsultation interactions in Hindi are provided to benchmark hybrid ASR-LLM systems. The dataset is designed to test the resolution of translation gaps, semantic drift, and phonetic errors by mapping colloquial symptom descriptions to SNOMED-CT and ICD-11 ontologies. It was authored by Aryan Raj Thakur and last updated on June 9, 2026.
Pidgin ASR Combined is a unified Nigerian Pidgin English speech-to-text dataset created by michaelodafe. It contains approximately 8.6 hours of audio across 4,278 clips from 10 source speakers, formatted as 16 kHz mono WAV files. The dataset was last updated on 2026-05-13 and was used to train a Whisper model that achieved a 21.37% word error rate.
1,200 code-switching utterances form a curated benchmark for evaluating commercial Automatic Speech Recognition systems. The dataset, created by Perle-ai, includes 300 samples each for four language pairs, such as Egyptian Arabic–English. It was last updated on May 21, 2026.
sWuggy is a spoken lexical-discrimination benchmark for evaluating spoken language models. Each item is a pair of a real word and a phonotactically matched pseudo-word, synthesized as audio. The dataset is hosted by the author 'coml' and was last updated on 2026-05-29.
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track includes full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema. The dataset was created by author Kukedlc and last updated on 2026-05 25.