Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,586 datasets
Music Extract Labeled Features is a dataset published on Kaggle. The dataset likely contains audio samples or tracks with associated feature labels. Metadata is minimal; actual content requires verification after download.
Multiple test sets for evaluating zero-shot text-to-speech models, introduced in associated research papers. The collection includes Chinese and English dialogue testsets, an English LibriSpeech testset, and a 24-language multilingual testset. The repository is maintained by k2-fsa and was last updated in April 2026.
ASR Models is a dataset hosted on Kaggle. The dataset's title suggests it contains resources related to Automatic Speech Recognition. Specific details about the data's size, origin, and structure are not provided in the available metadata.
June and July 2024 field data collected by the Alaska Division of Geological & Geophysical Surveys across an approximately 24,479 km2 area near Anaktuvuk Pass, Alaska. The dataset presents field station locations, observations, sediment sample descriptions, grain-size analyses, and links to photographs. This work was completed in support of a sand and gravel resource assessment for the Arctic Strategic Transportation and Resources (ASTAR) project.
Pittsburgh region meteorological data collected between July 2001 and November 2002 as part of the EPA's Particulate Matter Supersite Program. The dataset includes measurements of temperature, relative humidity, precipitation, wind speed and direction, UV intensity, and solar intensity from a central site and five satellite locations. It was produced by the NARSTO partnership and archived by NASA.
Cc100 Nepali TTS Shristi Encoded is a dataset for text-to-speech (TTS) applications in the Nepali language. The dataset was uploaded by user lilgoose777 to the Hugging Face platform and was last updated on May 31, 2026. Its specific content and structure are not detailed in the available metadata.
MinSpeech is a cleaned, multi-dialect Min-nan (Southern Min) speech dataset maintained as a private research fork by user 'scbz'. The dataset is intended for fine-tuning Automatic Speech Recognition and Speech-to-Text Translation models. It was last updated on the Hugging Face platform on April 20, 2026.
Synthetic-ASR-HI is a dataset hosted on Kaggle. The title suggests it contains synthetic audio data for Hindi speech recognition. The dataset's specific contents, size, and creation details are not provided in the available metadata.
An audio dataset likely intended for Automatic Speech Recognition (ASR) tasks, published on Kaggle. The dataset's specific content, size, and collection details are not provided in the available metadata. Further verification after download is required to confirm its scope and quality.
Palynology and paleoecology data from the Mattson Formation in northwest Canada, published by the Government of Yukon. The dataset was last updated in March 2026 and is available under a yk-oglyk license.
LibriSpeech Segment is an English read-speech corpus with phone-level time alignments generated by the Montreal Forced Aligner. The dataset is derived from the LibriSpeech corpus and is suitable for training and evaluating phone recognition and phonetic segmentation models. It was created by changelinglab and last updated on the platform in April 2026.
Indian Ascent Common Voice Dataset is a speech audio collection hosted on Kaggle. The dataset likely contains voice recordings, potentially for training speech recognition models. Specific details on size, contributors, and recording dates are not provided in the available metadata.
RTTS-datasets is a collection of data hosted on Kaggle. The dataset's specific content, size, and structure are not described in the available metadata. Its origin, creation date, and detailed composition require verification after download.
DELEGATE52 is a benchmark dataset for evaluating large language models on long-horizon delegated document editing across 52 professional document domains. The dataset was developed by Microsoft to study the readiness of AI systems for delegated workflows, where knowledge workers instruct LLMs to edit documents on their behalf over long sessions. It was last updated on 2026-04-20.
CSV_AFTER_JSON_DATASRT is a dataset hosted on Kaggle. The title suggests it may contain data that has been converted from JSON format to CSV. The dataset's specific content, size, and origin are not detailed in the available metadata.
The Armed Conflict Location & Event Data Project (ACLED) provides weekly aggregated counts of political violence, civilian-targeting, and demonstration events in Saint Kitts and Nevis. Organized by country-year and country-month intervals, the data enables longitudinal monitoring of conflict trends through early 2026.
Northwest Atlantic Ocean temperature-depth profiles were collected via expendable bathythermograph (XBT) and bathythermograph (BT) casts from the vessels PITTSBURGH and ST. LOUIS between June 1984 and June 1986. The data were processed into the NODC C116 format, which records temperature at non-uniform inflection points to define the profile curve. This dataset was created for the Frontal Air-Sea Interaction Experiment (FASINEX) with support from the Ships Of Opportunity Program (SOOP).
A dataset related to text-to-speech synthesis, specifically concerning runtime performance. It is hosted on the Kaggle platform. The dataset's specific content, size, and creation details are not provided in the available metadata.
2,740 audio recordings at 16 kHz form a dataset for traditional Chinese speech synthesis and recognition. Each entry includes the audio, its length, traditional Chinese text, and a corresponding normalized simplified Chinese text. The dataset, originally from ivanzhu109/zh-taiwan, is mirrored and formatted by lianghsun.
A professional-grade German text-to-speech training corpus created by author semidark and last updated on 2026-04-13. It combines high-quality human narration from the HUI Audio Corpus and LibriVox with synthetic augmentation to provide a legally safe alternative for training models like kokoro. The dataset is described as a work in progress.