Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,573 datasets
Data from two follow-up studies to Chandler and Pronin (2012) investigating the effects of movement on thought speed and subsequent risk-taking behavior. The dataset includes raw data from a Fitts' tapping task (FittsBart), a lower limb tapping task (BodyBart), and the Balloon Analogue Risk Task (BART) for measuring pumping behavior, along with PANAS measures. The 5.6 MB dataset was authored by Clare MacMahon and last updated on 2026-05-25.
rtxtd's Tts Dataset Guj Eng is a high-quality Text-to-Speech training dataset. It contains 120 audio samples totaling 60 minutes of speech, evenly split between Indian English and Gujarati languages. The dataset was last updated on 2026-06-17.
Cambodian cultural speech data comprising 134.6 hours of manually curated speech-text pairs in the Khmer language. The dataset was created by DDD-Cambodia using eight native speakers and was last updated in May 2026. Recordings average 8.54 seconds in length and include speaker metadata such as gender, age group, and origin city.
37,000 km² of Yukon Territory are underlain by potential coal-bearing rocks from Mississippian to Tertiary periods. This inventory, produced by the Government of Yukon, documents coal occurrences in seven distinct geological areas and includes a 1:2,000,000-scale map. The extent of deposits is largely unknown, as detailed examination has been limited.
A geological reconnaissance report describes the Toobally fault and surrounding rock formations in northern Toobally Lake, Yukon. The report identifies a newly proposed Toobally Formation diamictite estimated at 1800 m thick and an 850-m-thick basalt succession. It was published by the Government of Yukon and last updated in April 2026.
Dataset Strix is a unified audio dataset for the STRIX project by the Proteus Group student association. It contains 2-second mono clips at 16 kHz for classifying sounds in conflict zones. The dataset was created by Tairooonz and last updated on June 12, 2026.
A 2015 dataset from the Queensland government's State of the Environment reporting describes the composition of litter. It notes that cigarette butts are the most common type of litter, despite constituting a small volume. The dataset was published by the Department of Environment, Tourism, Science and Innovation and is available under a CC-BY-4.0 license.
NV-Bench is the first benchmark for evaluating nonverbal vocalizations in text-to-speech systems, grounded in a functional taxonomy. It was created by CharlesNi and last updated on June 12, 2026. The benchmark aims to provide standardized metrics and reliable ground truth references for this expressive component of speech synthesis.
22 healthy adults participated in a real-time fMRI neurofeedback experiment using a novel musical interface. The dataset includes pre- and post-session questionnaire results assessing mood and subjective experience, alongside neuroimaging data from a 50-minute MRI session. The research was authored by Alexandre Sayal and shared on figshare with a CC-BY-4.0 license.
Med-Dictate is an evaluation dataset released by Corti ApS alongside the Symphony for Speech Recognition white-paper. It contains medical notes dictated by Corti team members and one contractor, with their written consent, in English, French, and German. The dataset is built for benchmarking automatic speech recognition and related NLP systems on medical-domain audio.
Pitt Ads Dataset (PittAdsDB) contains advertising artifacts from the University of Pittsburgh. The dataset includes image annotations for actions, reasons, and sentiments, as indicated by the available JSON files. The dataset was uploaded by Mindykkyan and was last updated on June 20, 2026.
Med-Term is a fully synthetic evaluation dataset for automatic speech recognition released by Corti ApS. It contains medical notes dictated via text-to-speech in German, French, and English, with no real patient data or PHI. The dataset was released alongside the Symphony for Speech Recognition white-paper.
Nineteen participants with moderate to severe Alzheimer's disease from four nursing homes participated in a single-group intervention study. The data, published in 2026, includes assessments of social engagement, episodic memory, observed emotion, and verbal interactions collected at baseline, during nine workshops, post-intervention, and at a one-month follow-up. The dataset is a 102.8 KB PDF file containing the study's data sheet, authored by Mikael Genguelou.
Nineteen voluntary residents with moderate to severe Alzheimer's disease from four nursing homes participated in a single-group intervention study. The data likely contains assessments of social engagement, episodic memory, and observed emotions collected at baseline, three points during the intervention, post-intervention, and a one-month follow-up. The dataset was authored by Mikael Genguelou and last updated on April 22, 2026.
MagicHub provides a dataset of scripted speech recordings in the Henan dialect of Chinese. Audio files are recorded in a quiet indoor environment at 48 kHz and 16 bits. The dataset is released under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Approximately 187 hours of Arabic speech recordings and transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud. The dataset was created by oddadmix to support Arabic speech technology research and development. It was last updated on the platform in June 2026.
MagicHub's Magicdata-Dialect-Northeastern Chinese-TTS-Lite dataset provides scripted speech recordings in a Chinese dialect. Audio files are recorded in a quiet indoor environment at 48 kHz and 16 bits in WAV format. The dataset is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
A synthetic speech corpus for Turkish medical automatic speech recognition research. Clinical sentences were synthesized using Google Cloud Text-to-Speech Chirp 3 HD voices. The dataset was created by turkmedstt and was last updated on June 10, 2026.
Herney Andrés García‐Perdomo authored a protocol for the first systematic review and meta-analysis on the role of music in plastic surgery settings. The protocol describes the planned methodology for aggregating and analyzing existing research on this topic. Its publication status is indicated as Open Access (green).
Decoded transcripts and word-level error metrics from the Gemma 4 Unified models evaluated as automatic speech recognition systems. The dataset includes results for two model sizes, gemma4:e4b and gemma4:12b, on three standard English test sets. The evaluation was produced locally with ollama and includes a script for full reproducibility.