Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
A dataset for music genre classification, likely containing audio files or features for a machine learning challenge. It was published on the Kaggle platform. The specific collection method, size, and temporal coverage are not detailed in the available metadata.
Featuring conversational and phrasal speech training and test data for the Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript, provided by Microsoft and SpeechOcean.com for research purposes.
Speakers_xtts is a dataset hosted on Kaggle. Its title suggests it contains audio data related to speech synthesis, likely for text-to-speech applications. The dataset's specific content, scale, and origin are not detailed in the available metadata.
speakers_xttsv3 is a dataset hosted on Kaggle. The title suggests it contains audio samples for text-to-speech applications. The dataset's author, organization, and specific content details are unknown.
A Vietnamese text-to-speech dataset containing 1,805 paired audio recordings and text transcriptions for fine-tuning VieNeu-TTS models. The dataset was created by author 'quocs' and last updated on February 10, 2026. Audio files are in WAV format at 24kHz, mono, with 16-bit PCM encoding.
Packed with an index analyzing the impact of 109 independent music venue zones on 4,190 local businesses across the United States. It was created by Stanislas Renard to measure how these zones reinforce local economic resilience, with approximately 95% of surrounding businesses being locally owned. The data categorizes establishments by type and distinguishes between total business impact and specific local business impact.
Phase 0 UVR5/Demucs vocal + instrumental stems from the Music Foundry project. The dataset likely contains separated audio tracks for music source separation tasks. It is hosted on Kaggle, but details on size, format, and creation date are unspecified.
An audio classification dataset published on Kaggle. The dataset likely contains audio samples with associated labels for classification tasks. Specific details on size, source, and creation date are not provided in the available metadata.
A speech synthesis dataset for the Hausa language. It was published on Kaggle, but the author, organization, and creation date are unknown. The dataset's size, specific content, and structure are not detailed in the available metadata.
bn-bd-tts is a dataset hosted on Kaggle. The title suggests it contains data for Bengali text-to-speech synthesis, likely including audio recordings and corresponding text transcripts. Specific details on volume, creator, and update history are not provided in the available metadata.
EMOPIA is a dataset of 1,087 pop piano music clips from 387 songs, annotated with clip-level emotion labels by four dedicated annotators. It was created by researchers including Hsiao-Tzu Hung from Academia Sinica and presented at ISMIR in 2021. The dataset includes multi-modal data in audio and MIDI formats.
A collection of Khmer speech audio files and corresponding transcripts sourced from the Women's Media Centre of Cambodia (WMC) website. The dataset is prepared for machine learning tasks, with scripts provided to process audio and metadata into Parquet files. It was created by user 'vichetkao' and last updated on February 21, 2026.
A Kazakh speech audio dataset published on the Hugging Face platform by the organization ai4kazakh. The dataset was last updated on March 30, 2026. The specific content, size, and collection methodology are not detailed in the available metadata.
XTTS Checkpoint 3000 is a dataset published on Kaggle. The title suggests it contains a checkpoint for an XTTS (text-to-speech) model, likely used for speech synthesis tasks. The specific content, size, and origin of the checkpoint require verification after download.
Trail map locations for select preserved lands along the Massachusetts coast. The dataset is provided by the organization SCIOPS via the NASA Earthdata platform.
A geologic map characterizes the sea floor of Western Massachusetts Bay. It was constructed by the CEOS_EXTRA organization using sidescan-sonar imagery, photography, and sediment samples. The temporal coverage and specific data volume are not provided.
71 articles and 474 FAQs comprise this text corpus focused on UK live music. Published on Kaggle, the dataset likely contains blog posts and guides related to music events. The raw description indicates a total of 174,000 words across the collection.
Samples collected from Ace, Watts, Dingle, and Williams Lakes in Antarctica in 1994. Three to five 13 cm diameter ice cores were taken from each lake, sectioned at 20 cm intervals, and examined under a microscope at Davis station. No living cells were found, and as a result nothing was published.
Calendar years 384 to 239 before present (BP) of tree-ring width measurements from the Coutts Chemist Building site in Whangarei, New Zealand. This paleoclimatology dataset is part of the NOAA NCEI World Data Service for Paleoclimatology archive, contributed by the NOAA National Centers for Environmental Information.
5 college students spent 2 months annotating shrimps for use with the YOLO26 object detection model. The dataset is designed for computer vision tasks related to counting and detection. Its specific scale and annotation methodology are detailed in the provided description.