Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
Seventeen years of ocean current data were collected by the U.S. Geological Survey using current meters deployed in New England coastal waters. The collection includes U and V velocities in cm/s, rotor speeds, current directions in degrees, and water temperatures. Data collection spanned from May 1975 to February 1992.
Expert scores and audio features for music assessment provide a structured evaluation of musical performances. The dataset likely contains quantitative metrics derived from audio recordings alongside subjective ratings from experts. Its origin and scale are unspecified.
A cleaned, metadata-rich Shona speech dataset prepared through a reproducible data engineering pipeline. The dataset is derived from the google/WaxalNLP source, specifically the sna_asr subset, and was last updated on March 20, 2026. It is intended as a general-purpose standard corpus for downstream tasks.
A multimodal dataset combining audio and text data for music genre classification. It likely contains audio features from the GTZAN benchmark dataset paired with corresponding song lyrics. The dataset is published on Kaggle, but its specific creation date and author are unknown.
Dataset_train_xtts is a dataset for training text-to-speech models, published on Kaggle. The dataset's specific content, size, and origin are not detailed in the provided metadata. Further details about the data's collection method, author, and temporal coverage are unavailable.
10,000+ hours of interview audio and video sourced for AI training. The data is described as ethically sourced. The dataset is hosted on Kaggle, but details about the author, organization, and specific collection dates are unknown.
Between 100,000 and 1,000,000 Spanish audio segments and transcriptions derived from LibriVox audiobooks. Created by Cnam-LMSSC and updated in March 2026, it extends the Multilingual LibriSpeech (MLS) corpus with machine-generated phonetic transcriptions.
A Kaggle dataset titled 'data_tts_32'. The dataset likely contains audio files and associated text for text-to-speech synthesis tasks. Its specific content, size, and origin are not detailed in the provided metadata.
An invoice dataset published on Kaggle. The dataset likely contains records related to business transactions and billing. Specific details regarding its size, origin, and time period are not provided in the available metadata.
UK Live Music Booking Rates 2026 (May) contains 3,847 booking rates for live music in the UK. The data is organized by city, band size, and event type. The dataset was sourced from Kaggle, but the author, organization, and specific collection method are unknown.
32,267 audio samples totaling 103.18 hours of Vietnamese speech, curated for automatic speech recognition. The dataset, created by thanhnew2001, was last updated in February 2026. It is structured into 29,041 training and 3,226 development samples.
XTTSv2Audios is a dataset of audio files likely associated with the XTTSv2 text-to-speech model. The dataset is hosted on Kaggle, but its specific contents, size, and creation details are not provided in the available metadata. Further details such as the number of samples, speaker diversity, and recording conditions require verification after download.
Tts Male 70H is a text-to-speech dataset published on HuggingFace by user vfdanil. The title suggests it contains audio samples of a male voice, likely for speech synthesis tasks. The dataset was last updated on April 22, 2026.
Kaggle hosts a dataset titled 'Noise reduction'. The dataset's content, size, and specific source are not detailed in the provided metadata. Its last update date and licensing information are also unknown.
Ayf3 published the Numberblocks One Voice Dataset on Hugging Face in April 2026. The dataset likely contains audio recordings related to the Numberblocks children's media franchise. Its specific content, size, and structure require verification after download.
Sarah Lyons Watts's book presents a psychological portrait of the 26th U.S. president, Theodore Roosevelt. The work analyzes his personal obsession with masculinity and its influence on national politics, as noted by contemporary figures like Woodrow Wilson. It is a textual analysis of Roosevelt's legacy, sourced from the paperswithcode platform.
Coastal Massachusetts dive sites include reefs, wrecks, jetties, and breakwaters popular for SCUBA diving. Data points were compiled by the Massachusetts Office of Coastal Zone Management from the Board of Underwater Archaeological Resources and dive club listings. The Massachusetts Office of Coastal Zone Management updated this layer on July 2, 2007.
Point locations document federal dredge projects by the US Army Corps of Engineers along the Massachusetts coastline. The data is historical, with records up to 16 December 1998. The dataset was compiled by the organization SCIOPS.
Polygonal extents document federal dredging projects by the US Army Corps of Engineers along the Massachusetts marine coastline. The dataset includes navigational channels, anchorages, harbors, beaches, and dikes, with records historical to December 16, 1998. It was compiled by the organization SCIOPS.
Geospatial arcs represent commercial, charter, and recreational boating uses within the Massachusetts Coastal Zone. The data delineates three distinct activity subtypes. It was compiled by SCIOPS from expert workshops held in Boston and Waquoit in June 2005.