Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
A multi-speaker clinical speech corpus containing nursing handover statements. It is designed for research in Automatic Speech Recognition and speech-driven clinical documentation, featuring speakers with different English accents.
PersianPunc is a large-scale dataset for Persian punctuation restoration, containing 17 million token-level sequence labeling samples aggregated from 6 source corpora. It was created by MohammadJRanjbar and accepted at the EACL 2026 SilkRoad NLP Workshop.
tw-hokkien-seed-text is a dataset of approximately 3 million full-character Taiwanese Hokkien sentences designed for training text-to-speech (TTS) and automatic speech recognition (ASR) models. The dataset was created by lianghsun and was last updated on March 20, 2026. Each sentence is 50โ80 characters long, corresponding to a speech duration of 10โ15 seconds, and is written exclusively in Chinese characters to preserve authentic Taiwanese Hokkien vocabulary and syntax.
IndicTTS-p2 is a dataset for text-to-speech synthesis, likely containing audio recordings and corresponding text transcripts. It is hosted on Kaggle, but the author, organization, and specific data characteristics are not provided. The dataset's size, format, and exact language coverage are unknown from the available metadata.
Genshin Matcha TTS is a dataset hosted on Kaggle. The title suggests it contains audio data for text-to-speech synthesis, likely related to the 'Genshin' context. No further metadata on size, source, or creation date is available.
Abjad-Kids is an Arabic speech classification dataset containing spoken recordings of the Arabic alphabet, numbers, and colors from multiple child speakers. It supports research in automatic speech recognition and educational technology for Arabic-speaking children. The dataset was created by Aziz-snoubra and was last updated on March 14, 2026.
A speech dataset intended as an example for training a text-to-speech fine-tuning platform. It contains audio files with associated transcripts and speaker identifiers, with missing transcripts generated automatically by the Whisper-large v3 model. The dataset was created by mgrei and was last updated on April 12, 2026.
13 primary spoken languages, including English, Spanish, Mandarin, and Hmong, are tracked for individuals who enrolled in a Covered California Qualified Health Plan. The data originates from the California Healthcare Eligibility, Enrollment and Retention System (CalHEERS) and is part of public reporting requirements. Enrollment counts are reported by period for individuals who paid their first premium.
A sample dataset of high-fidelity, ethically sourced conversational audio data. The description indicates it is intended for voice cloning applications. The dataset's size, specific source, and temporal coverage are unknown.
Between 1875 and 1895, the prevalence of double-entry bookkeeping among Massachusetts corporations surged from 60% to over 96%. This dataset supports a quantitative analysis of accounting innovation, tracking the adoption of depreciation and its correlation with firm survival. It includes balance statement data for corporations and supplementary citation counts from the Accountants' Index.
James D. Tucker's fdasrvf package implements the square-root velocity framework for elastic functional data analysis. The method, based on research by Srivastava et al. (2011) and Tucker et al. (2014), performs alignment, PCA, and modeling of multidimensional and unidimensional functions. It is sourced from the paperswithcode platform.
Approximately 1,700 musical pieces in MP3 format, sourced from NetEase music. The audio clips are 270 to 300 seconds long and sampled at 22,050 Hz. The dataset was created by ccmusic-database and last updated on 2026-02-27.
A speech dataset covering multiple regional dialects of the Bangla language, intended for automatic speech recognition tasks. The dataset is hosted on Kaggle, but details on its size, collection method, and creator are unspecified. Its primary focus is on capturing linguistic diversity within the Bengali-speaking regions.
A high-quality, speaker-paired subset of the LAION Emolia dataset, created by TTS-AGI and last updated on March 9,ๆไปฌๅ็ฐ 2026. Each sample includes a target and a reference utterance from the same speaker, filtered for quality using a DNSMOS score threshold of 3.0.
CMI Pref Pseudo contains 56,000 music generations from 23 models and 165,000 pairwise comparisons for preference modeling research. The dataset was created by HaiwenXia and last updated on March 3, 2026. Prompts are compositional, including text, optional lyrics, and reference audio.
IMSLP MIDI Dataset contains MIDI files and associated metadata crawled from the International Music Score Library Project on July 21-22, 2024. The dataset includes fields such as composer, year, era, style, key, and license, along with raw MIDI bytes and serialized mido objects. It was created by TiMauzi and is available under a CC-BY-SA-4.0 license.
722 seed utterances and 32,506 Common Voice samples were used to generate this Taiwanese Hokkien (Min Nan) speech dataset via the CosyVoice3 model. The dataset includes audio files, corresponding text, and speaker metadata. It was created by lianghsun and last updated on March 19, 2026.
WeWe Pidgin TTS Dataset is a speech synthesis dataset published on Kaggle. The dataset likely contains audio recordings and corresponding text transcriptions for text-to-speech applications. Its specific size, creation details, and update history are not provided in the available metadata.
ToneWebinars Balalaika is a 248.9-hour Russian speech corpus curated from podcasts by the MTUCI lab260 team. Released in early 2026, the dataset was processed using the BALALAIKA pipeline to provide high-quality audio for generative speech tasks. It serves as a refined version of the original ToneWebinars source, specifically filtered for speech synthesis and recognition.
TWB Voice Kanuri TTS 1.0 Sample Set is a high-quality text-to-speech corpus containing read speech data in Kanuri. It was recorded by a single female speaker under acoustically optimal conditions and represents 10% of the complete dataset collected by CLEAR Global (formerly Translators without Borders). The dataset page was last updated on 2026-02-23.