Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,586 datasets
Seed Tts Eval Arrow is a dataset for evaluating text-to-speech systems, published on HuggingFace by zhaochenyang20. The dataset was last updated on 2026-05-22. Its specific content and scale require verification after download.
Dari Wavs is an audio dataset created by Sanji27. The description suggests the dataset could be expanded in size and include transcripts ready for automatic speech recognition (ASR). The dataset was last updated on May 17, 2026.
Anonymized streaming logs and composition statements for royalty auditing. The dataset appears to be a benchmark for detecting royalty leakage in the music industry. Its specific temporal coverage, size, and creator are not detailed in the provided metadata.
wild_asr_haid is a dataset hosted on Kaggle. The title suggests it contains audio data, likely for automatic speech recognition (ASR) tasks. Its specific content, size, and origin require verification after download.
Voicebench Ja contains 4 subsets created by applying speech synthesis to samples from three Japanese text benchmarks: Elyza-tasks-100, M-IFEval, and JamC-QA. The dataset was constructed by SB Intuitions using their internal TTS model and JVS corpus audio prompts to quantitatively evaluate performance gaps between audio and text inputs for language models. It was last updated on March 30, 2026.
A test audio dataset for the ADLIB language-aware ASR benchmark framework for Japanese. It contains 247 test cases with audio from 3 speakers, focusing on the DevTerm (software development terminology) domain. Reference transcripts and term annotations are provided in a separate JSONL file within the project's GitHub repository.
529 audio segments totaling 46 minutes provide speech data for Turkana, an Eastern Nilotic language with roughly 1 million speakers in Kenya. The dataset was created by Speedykom using Bible narratives from the Global Recordings Network, segmented via silence detection. Transcripts were auto-generated using the facebook/mms-1b-all model with a Teso adapter.
A text-to-speech dataset published on HuggingFace by author thach124. The dataset was last updated on 2026-05-28 10:23:30. Its specific content and scale are unknown from the provided metadata.
NASA's CYGNSS constellation provides calibrated Delay Doppler Maps (DDMs) measuring ocean surface scattering. The dataset contains daily files from up to 8 spacecraft, with a typical latency of 6 days from measurement. Version 3.1, produced by POCLOUD, supersedes Version 3.0 with improved antenna gain calibration and quantization correction.
Structured metadata for Greek laΓ―ko music tracks, intended for research and machine learning. The dataset includes fields for emotion, era, and genre but does not contain audio files. It was created by author christosfouk and was last updated on 2026-04-16.
Librispeech-PC 44kHz Opus replaces the original Librispeech PC audio with higher-quality source material encoded as Opus at 64 kbps. Sampling rates are increased from 16kHz up to 48kHz, depending on the source. The dataset was created by mythicinfinity and last updated on March 28, 2026.
A time-series dataset of the Gold (XAU) to US Dollar (USD) exchange rate, likely recorded at 3-minute intervals. The dataset is published on Kaggle, but its author, size, and specific time range are unknown. The raw description suggests it contains price data for the forex pair XAU/USD.
A SQLite database contains user votes and feedback from TTS Arena, a platform for comparing text-to-speech models. The dataset was created by Pendrokar and last updated in April 2026. It is designed to help developers identify model faults through community evaluation.
A 2016 geospatial dataset from NOAA characterizing macroalgae beds for oil spill sensitivity planning in Massachusetts and Rhode Island. Vector points represent vegetation beds, with associated tables containing species-specific abundance, seasonality, and life history information. The data is part of a larger Environmental Sensitivity Index (ESI) effort to map coastal resources.
National Oceanic and Atmospheric Administration (NOAA) data for Massachusetts and Rhode Island contains sensitive biological resource data for benthic species. Vector polygons represent submerged aquatic vegetation and macroalgae, with associated tables for species abundance, seasonality, and life history. This data is part of the Environmental Sensitivity Index (ESI) characterizing coastal environments by their sensitivity to oil spills.
A speech recognition dataset sourced from YouTube, likely containing audio and corresponding transcriptions. It was published by user 'veziriii' on Hugging Face and was last updated on May 23, 2026. The specific content and scale require verification after download.
SMAPVEX19-22 field campaign collected daily mosaicked UAVSAR images at three polarization configurations from April to July 2022 near Petersham, Massachusetts. The terrain-flattened gamma-corrected radar data targets forested land cover to validate satellite-derived soil moisture estimates. This dataset supports the Soil Moisture Active Passive Validation Experiment's goal of improving remote sensing accuracy in vegetated areas.
16,017 audio samples filtered from a larger 616-hour speech dataset to contain only ElevenLabs Scribe v1 audio events. The dataset, created by TTS-AGI, focuses on vocal bursts and background sounds with unified annotation formatting. It was last updated on March 28, 2026.
10 audio clips serve as a stage 0.5 acceptance check for the EmiratiTTS project, fine-tuned for Emirati Arabic. The clips were created by Alqayed2024 to verify the data, tokenizer, and pipeline wiring before a full fine-tuning run. The dataset page was last updated on April 12, 2026.
Librispeech_manifests likely contains metadata files for the LibriSpeech corpus, a widely used benchmark in automatic speech recognition. The dataset is published on Kaggle, but its specific contents and scale are not detailed in the available metadata. Columns and sample data are unknown, requiring verification after download to confirm the exact structure and utility.