Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,587 datasets
Maine's coastline from Cutts Island to Prouts Neck is covered by ortho-rectified mosaic tiles. The National Oceanic and Atmospheric Administration (NOAA) produced this data from imagery acquired June 5-7, 2011, using an Applanix Digital Sensor System (DSS). The final mosaic is derived from higher-resolution original aerial photographs.
NOAA's Integrated Ocean and Coastal Mapping initiative produced this ortho-rectified image mosaic. The source aerial photographs were captured with an Applanix Digital Sensor System between June and September 2011. The final mosaic covers ports in the Cape Cod region of Massachusetts.
Coastal Maine from Cutts Island to Prouts Neck is covered by ortho-rectified mosaic tiles from the NOAA Integrated Ocean and Coastal Mapping initiative. The source imagery was acquired on June 7, 2011, using an Applanix Digital Sensor System (DSS). The final ortho-rectified product is derived from higher-resolution original images.
NOAA's 2011 ortho-rectified mosaic tiles were created under the Integrated Ocean and Coastal Mapping initiative. The source imagery was acquired from June to September 2011 using an Applanix Digital Sensor System. The original aerial photographs were captured at a higher resolution than the final mosaic product.
NOAA NGS ortho-rectified mosaic tiles created from imagery acquired between August 10 and October 21, 2009. The National Oceanic and Atmospheric Administration produced this data through its Integrated Ocean and Coastal Mapping initiative using an Applanix Digital Sensor System. The source imagery was acquired at a higher resolution than the final mosaic product.
Ortho-rectified mosaic tiles were created from aerial imagery acquired on June 7, 2011, using an Applanix Digital Sensor System (DSS). This data product is part of the NOAA Integrated Ocean and Coastal Mapping initiative, covering the Maine coastline from Cutts Island to Prouts Neck. The source images were acquired at a higher resolution than the final ortho-rectified mosaic.
Ortho-rectified mosaic tiles created by NOAA's Integrated Ocean and Coastal Mapping initiative. The source aerial imagery was captured from June 5 to June 7, 2011, using an Applanix Digital Sensor System. The final product is derived from higher-resolution original images.
New Bedford, Massachusetts is covered by ortho-rectified mosaic tiles produced by the NOAA Integrated Ocean and Coastal Mapping initiative. The source imagery was acquired on October 5, 2011, using an Applanix Digital Sensor System aircraft. The final mosaic is derived from higher-resolution original images.
A unified Danish speech recognition dataset combines approximately 3.5 million audio samples from seven distinct sources, totaling roughly 16,000 hours of speech. The collection includes European and Danish Parliament recordings, read-aloud and conversational speech, broadcast media, and crowd-sourced samples. It was created by syvai and last updated on the Hugging Face platform in April 2026.
Google Waxal ASR Challenge data likely contains audio recordings and transcriptions for automatic speech recognition benchmarking. The dataset is hosted on Kaggle, a platform for data science competitions. Its specific size, collection method, and time range are not detailed in the available metadata.
DMSP OLS satellite data provides visible and infrared imagery for monitoring global cloud distribution and cloud top temperatures twice daily. The archive includes low-resolution global and high-resolution regional imagery from a 3,000 km scan, alongside satellite ephemeris and solar/lunar information. Data is sourced from the DMSP Operational Linescan System instruments and archived by NOAA NCEI.
A bilateral social security agreement and administrative arrangement between Canada and the Federation of Saint Kitts and Nevis. The agreement coordinates the two countries' social security systems for individuals who have lived or worked in both jurisdictions. It was published by Global Affairs Canada and is archived as of February 2026, indicating it is out of date and for research purposes only.
A speech dataset likely intended for text-to-speech research, hosted on Kaggle. The dataset's author, organization, and specific content details are not provided in the metadata. Its original creation date and update history are unknown.
Visual Wake Words - Corrupted (VWW-C) is a dataset derived from the Visual Wake Words benchmark, likely containing images with synthetic corruptions. Published on Kaggle, its specific size, license, and authorship details are not provided. Columns and sample data are unknown, suggesting metadata is minimal.
TaigiSpeech is a spoken language understanding dataset containing over 3,000 Taiwanese speech utterances from 21 speakers. Each utterance is labeled with one of 8 intent classes, designed for elder-care and smart-home voice command scenarios to support research in a low-resource language.
A 5.5 KB tabular dataset documents the surface preservation condition of Tritia cf. gibbosula mollusk shells from archaeological unit US 8 at El Mnasra cave. The dataset, authored by Emilie Campmas and shared under CC BY 4.0, provides a taphonomic record for paleontological and archaeological analysis.
Processed audio files for Tunisian Arabic Automatic Speech Recognition (ASR). The dataset is hosted on Kaggle, but its size, creation date, and author are unknown. The title suggests it contains audio data that has been processed for use in speech recognition tasks.
ASR_uaspeechdata is a dataset published on Kaggle. The title suggests it contains audio data likely intended for training or evaluating automatic speech recognition systems. The dataset's specific content, size, and origin are not detailed in the available metadata.
Slakh2100 is a large-scale dataset containing 2,100 automatically mixed music tracks with isolated instrument stems and aligned MIDI data. Created by Manilow et al. in 2019 at Northwestern University, it is designed for music information retrieval and source separation research. The dataset is hosted by schism-audio on Hugging Face.
104,478 fully synthetic duplex conversations provide 2,133 hours of 16kHz audio for training real-time spoken dialogue models. The dataset was created by author mailong225 for the RelayS2S hybrid architecture, converting text dialogues to speech. It was last updated on March 25, 2026.