Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,589 datasets
VOXCeleb is a dataset of speech and video clips featuring celebrities. It is hosted on the Kaggle platform. The specific size, collection method, and time range are not detailed in the provided metadata.
TTS13G is a dataset hosted on Kaggle. Its title suggests a focus on text-to-speech synthesis, likely containing audio samples and corresponding text transcripts. The dataset's specific size, origin, and detailed contents are not described in the provided metadata.
Point-based GIS data showing the locations of marinas, yacht clubs, boat yards, and related facilities along the Massachusetts coast. The data were compiled in 2007 from public lists, databases, and visual inspection of orthoimagery by the organization SCIOPS. All data are represented as points with associated attribute data and include facilities defined as catering to recreational yachtspersons.
June 2006 GIS data showing potential or developing locations for desalination plants along the Massachusetts coast. The data are preliminary and speculative, compiled from public media reports, state/federal regulatory filings, local meeting proceedings, and private contractor studies. The dataset was aggregated by SCIOPS and is hosted on the nasa_earthdata platform.
Massachusetts Bay hosts a geospatial data layer detailing a proposed 16-mile, 24-inch diameter natural gas pipeline lateral. The dataset, created by SCIOPS, represents the project layout as of September 27, 2005, based on surveys using DGPS, multibeam, sidescan, and diver inspections. It was surveyed according to US Army Corps of Engineers standards (EM 1110-2-1003).
A model artifact for sequential recommendation, published on Kaggle. The specific data format, size, and creation details are not provided in the metadata. The content and structure require verification after download.
A collection of characteristic FTIR peaks (cm⁻¹) for cigarette butts and beach sand, categorized by beach use zones. The data is provided in an XLSX file of 16.8 KB, authored by Claudia Díaz-Mendoza and last updated in March 2026.
A geospatial line feature representing the western portion of a 46kV electric supply cable from Hyannis, Cape Cod to Nantucket Harbor. The data was created by SCIOPS using geographic coordinates from a National Grid drawing dated February 2005, incorporating marine survey data from 1994.
Geospatial data from 1994 and 1996 shows the location of a 46kV submarine electric supply cable between Harwich Port, Cape Cod and Nantucket Harbor, Massachusetts. The line feature was created by SCIOPS using coordinates from a 1996 National Grid drawing and incorporates data from a marine survey conducted in 1994.
Voicedesign3 is a Vietnamese text-to-speech dataset created by ShiniChien. The dataset is synthesized, meaning the audio was likely generated by a TTS model rather than recorded from human speakers. It was last updated on HuggingFace on April 24, 2026.
VoxCeleb1 is an audio-visual dataset sourced from YouTube videos. It is published on the Kaggle platform. The dataset's specific size, creation date, and author are not detailed in the provided metadata.
CrispASR Kaggle CUDA Build is a dataset hosted on Kaggle. Its title suggests it relates to CUDA build tools, likely for the CrispASR speech recognition project. The dataset's specific content, size, and structure are not detailed in the available metadata.
Mono Segments contains over 310,000 multi-instrumental MIDI files selected from the Discover MIDI Dataset. The dataset is enriched with lead monophonic melodies and high-precision structural segment labels, created by author asigalov61.
STOMA is a multi-speaker Greek speech corpus containing approximately 23 hours of studio-recorded read speech. It features audio from six native speakers (three male and three female), captured under controlled studio conditions to ensure high signal quality.
Moroccan Darija ASR Dataset Split is a speech corpus for Automatic Speech Recognition, published on the Hugging Face platform by mohamedmou. The dataset was last updated on May 1, 2026, but its specific size, content, and collection methodology are not detailed in the available metadata.
I music is a dataset hosted on Kaggle. Its specific content and scope are not detailed in the available metadata. The dataset's origin, size, and creation date are unknown.
Results from the SASRec v1 model applied to Spotify data. The dataset is hosted on Kaggle. The specific content, size, and creation details are not provided in the metadata.
Uyghur language speech recordings for natural language processing tasks. The dataset contains 2,157 audio files in MP3 format, totaling 3.03 GB, created by user 'anke01' and last updated on February 26, 2026.
CoversBR is a large audio database focused predominantly on Brazilian music for cover and live song identification tasks. It comprises metadata and extracted features from 102,298 songs, organized into 26,366 cover groups, totaling approximately 7,070 hours of audio. The dataset is provided by Dirceu G Silva via AWS Open Data, but the original audio files are not included due to copyright restrictions.
MusicX is a dataset hosted on Kaggle. Its specific content and scale are not detailed in the provided metadata. The dataset likely contains audio data or features related to music, based on its title.