Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
High-resolution seismic-reflection surveys map the stratigraphy of the nearshore areas from Chatham to Provincetown, Massachusetts. The U.S. Geological Survey Woods Hole Field Center conducted this investigation to correlate geologic units between the nearshore and onshore. The data defines the Quaternary geologic framework of outer Cape Cod.
A 2006 data set provides bathymetric contours for the Gulf of Maine and New England Shelf. The U.S. Geological Survey constructed it for geologic framework studies. It was reprojected into the NAD83 Massachusetts State Plane coordinate system by the Massachusetts Office of Coastal Zone Management.
Monitoring data tracks the environmental effects of secondary-treated sewage effluent discharged from a 9.5-mile outfall tunnel into Massachusetts Bay. The Environmental Quality Department (Enquad) collects this data to ensure compliance with an NPDES discharge permit for 43 communities. The dataset covers water quality in Massachusetts Bay, Boston Harbor, and Cape Cod Bay.
Massachusetts Bay hosts the as-built location of the Hubline, a 29.5-mile natural gas pipeline constructed primarily offshore between Beverly and Weymouth. The dataset was created by SCIOPS, representing the pipeline's surveyed bottom position. The route traverses 11 coastal communities including Salem, Boston, and Quincy.
A dataset for Afaan Oromo text-to-speech synthesis, published on Kaggle. The dataset likely contains paired text and audio samples for training and evaluating speech synthesis models. Specific details on size, format, and collection methodology are not provided in the available metadata.
MLAAD English provides audio samples for evaluating text-to-speech models. The dataset likely contains five audio clips generated by each of several TTS systems. It is hosted on Kaggle, but the specific creator and update date are unknown.
1982 aerial photography of penguin colonies on islands approximately 12km northeast of Brattstrand Bluff, Antarctica, was digitized into DXF files and later georeferenced into a shapefile. The dataset includes digitized colony boundaries and four supporting photographs from 2009. Work was contributed by Eric Woehler, John Cox, Tom Velthuis, and Ursula Harris of the Australian Antarctic Data Centre.
NVIDIA's Granary dataset provides approximately 1 million hours of high-quality speech data across 25 European languages for speech recognition and translation. Released in 2026, it consolidates multiple sources into a unified framework to support low-resource language modeling. The collection is designed for both Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) tasks.
PICO-8 Games Dataset contains 10,967 game cartridges scraped from the Lexaloffle BBS. Each cartridge is decomposed into Lua source code, pixel-art spritesheets, tile maps, sound effects, music patterns, and metadata. The dataset was created by Fraser and includes label screenshots from the top 48 games by star count.
LibriVAD is a large-scale, noise-augmented dataset for Voice Activity Detection (VAD) generated from the LibriSpeech corpus. The dataset was created by LibriVAD and was last updated on March 17, 2026. It is designed for training and evaluating VAD models in noisy environments.
Vivoice Relabeled is a speech dataset derived from the original capleaf/viVoice collection. The dataset has been processed using the Qwen/Qwen3-ASR-1.7B model to update audio-text labels, retaining samples with a Word Error Rate below 15%. It was uploaded by author JayLL13 to Hugging Face in March 2026.
CMI-Pref provides between 1,000 and 10,000 human preference comparisons for multimodal music generation, published by HaiwenXia in 2026. Each record captures a human vote comparing two generated audio samples based on musicality, alignment, and confidence.
Naija-ASR-Corpus v1.0 (NAC-v1.0) is a foundational speech dataset for Nigerian Pidgin (Naija, PCM). It was created by the NAC Team, who processed long-form recordings from the Universal Dependencies Naija Spoken Corpus into sentence-level audio-text pairs suitable for ASR training. The dataset was last updated on March 16, 2026.
IndicTTS-p1 is a dataset for text-to-speech synthesis, published on Kaggle. The title suggests it contains data for Indic languages, which likely includes audio recordings and corresponding text transcripts. The dataset's specific size, languages, and collection details are not provided in the available metadata.
IndicTTS-p3 is a text-to-speech dataset likely containing audio samples and corresponding text transcripts for one or more Indic languages. It is hosted on Kaggle, but the author, organization, and specific data characteristics are not provided. The dataset's size, format, and exact contents require verification after download.
33,228 synthetic audio clips for Taiwanese Hokkien text-to-speech, generated using the Qwen3-TTS-1.7B-Base model with voice cloning. The dataset was created by lianghsun and last updated in March 2026.
Encompassing 2,700 pairs of text-to-speech audio renderings with 15 human preference annotations per pair. Produced by datapointai and updated in March 2026, it provides comparative naturalness ratings for audio generated from identical text prompts. The collection totals 40,500 individual human judgments to support high-confidence audio quality evaluation.
NileTTS provides 38.1 hours of transcribed Egyptian Arabic speech across 9,521 utterances, published by KickItLikeShika in February 2026. The collection is segmented into specific domains, including over 21 hours dedicated to sales and customer service interactions.
TTS Human Preferences (Medium) is a dataset for text-to-speech audio quality evaluation. It contains 2,000 rows, each with two TTS audio renderings and 15 human preference annotations, totaling 30,000 annotations. The dataset was created by datapointai and last updated in March 2026.
Ttsmodels is a dataset published on HuggingFace by author phongluong197. The dataset was last updated on 2026-05-07 07:59:33. Its specific content and scale are not detailed in the available metadata.