Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
A speech corpus for the Urdu language, published on Kaggle. The dataset likely contains audio recordings paired with corresponding text transcripts for training text-to-speech systems. Specific details on size, collection method, and contributors are not provided in the available metadata.
An audio dataset of Hindi speech, published on the Kaggle platform. The dataset likely contains audio files of spoken Hindi, which can be used for training and evaluating speech processing models. Specific details on the number of recordings, speakers, recording conditions, and collection methodology are not provided in the available metadata.
Georgian language audio recordings for text-to-speech synthesis, published on Kaggle. The dataset's size, collection method, and specific content are not detailed in the available metadata. Further details regarding the number of samples, recording quality, and speaker demographics require verification after download.
Emotion-Aware Music Sentiment Dataset provides multimodal audio features and contextual metadata for emotion-based music AI. The dataset originates from Kaggle, though specific details on volume, authorship, and recency are unavailable.
Satellite imagery data covers the Caribbean nation of Saint Kitts and Nevis. The dataset is provided by Techsalerator and hosted on Kaggle. Specific details on data volume, collection date, and resolution are not provided in the input.
Archive of Our Own (AO3) data related to music and bands, collected via web scraping. The dataset's size, row count, and specific attributes are unknown. The author, organization, and last update date are also unspecified.
Synthetic German medical speech data targets neuro-oncology and neurology terminology for fine-tuning ASR models. The dataset was generated by NeurologyAI using Resemble AI Chatterbox TTS for voice and Qwen/Qwen3-30B-A3B for text. It was last updated on January 15, 2026.
A curated collection of historical audio recordings sourced from the Library of Congress Citizen DJ collections. The dataset is designed for open research, audio analysis, music information retrieval, remixing, and AI/ML experimentation. This mini release is intended as a lightweight subset for testing pipelines, educational use, and small-scale experiments.
AsramaGH is a dataset published on Kaggle. Its title suggests it may contain information related to housing or community structures. The specific content, size, and origin are unknown.
AsramaGH likely contains data related to housing or community structures. The dataset is published on Kaggle, but its creator, size, and specific content are unknown. Its last update date is also unknown.
Child Trends provides data on student engagement in arts education. The dataset's specific variables, size, and temporal coverage are not detailed in the available metadata. The original source is listed as paperswithcode, a platform for machine learning resources.
A processed corpus for Urdu text-to-speech (TTS) applications, published on Kaggle. The dataset likely contains audio recordings and corresponding text transcriptions. Specific details on size, source, and processing methods are not provided in the available metadata.
A dataset titled 'dataindextts3' is hosted on Kaggle. The title suggests it contains data related to text-to-speech (TTS) synthesis. No further metadata, such as author, size, or sample details, is provided.
DataIndexTTS4 is a dataset published on Kaggle. Its title suggests it is related to text-to-speech (TTS) technology. The dataset's specific content, size, and origin require verification after download.
30 Musical Instruments is a dataset hosted on Kaggle. The title suggests it contains audio samples or information related to a collection of thirty different musical instruments. The dataset's specific content, size, and origin are not detailed in the provided metadata.
Urdu TTS Corpus Processed - 5 is a dataset for text-to-speech applications, published on Kaggle. The title suggests it contains processed audio and corresponding text data for the Urdu language. The specific content, scale, and creation details require verification after download.
VoxCeleb is an audio dataset hosted on Hugging Face. The dataset was uploaded by author N02N9 and was last updated on 2026-02-24. Its specific content, scale, and collection method are not detailed in the provided metadata.
17,476 preprocessed Urdu speech samples from Mozilla Common Voice, split into training, validation, and test sets. The dataset is processed for Whisper models, with audio resampled to 16kHz. It was uploaded by khawajaaliarshad and last updated on 2025-12-27.
The Journal of the Musical Arts in Africa is a source of academic publications. It likely contains scholarly articles and research papers on music and related arts from an African context. The dataset is aggregated from the paperswithcode platform.
modern-tts-dataset is a dataset for text-to-speech (TTS) research, published on Kaggle. The dataset likely contains audio recordings paired with corresponding text transcripts. Specific details on size, source, and creation date are not provided in the available metadata.