Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,602 datasets
An assessment of Pittsburgh's One Vision One Life violence prevention strategy authored by Jeremy M. Wilson. The report likely contains data on program implementation, operations, and impact, including community-building, conflict intervention, and mediation. It also includes comparisons with other cities and lessons learned.
The dataset likely contains historical analysis of rock music's role in Cold War geopolitics. It appears to be sourced from a research paper discussing containment policy and Western support for Yugoslavia. The specific data volume and structure are unknown.
A historical text traces the emergence and conflict of rock music culture in Eastern Europe and the Soviet Union from 1954 to the present. It covers the 30-year conflict between rock fans and the Communist Party, including events in Prague in 1968 and Poland in 1981. The source is a book titled 'Rock Around the Bloc', but the specific dataset format and structure are unknown.
my_asr_dataset_v2 is a dataset for automatic speech recognition, published on Kaggle. The dataset's specific size, collection method, and temporal coverage are not detailed in the available metadata. Its content and structure require verification after download.
A dataset titled 'MoviesTextSASRec' published on Kaggle. The title suggests it likely contains text data related to movies, potentially for use with sequential recommendation models like SASRec. The dataset's author, organization, size, and specific content are unknown.
ASRDemo is a dataset published on Kaggle. Its title suggests it contains audio data for speech recognition demonstration purposes. The dataset's specific size, format, and content details are unknown.
SeniorTalk provides 10,000 to 100,000 Mandarin Chinese speech records from individuals aged 75 to 85, produced by BAAI in 2025. It includes audio and text modalities to facilitate research in automatic speech recognition and speaker verification for the super-aged population.
An audio dataset titled 'my_asr_dataset' is hosted on Kaggle. The dataset's content, size, and specific characteristics are not detailed in the provided metadata. Its creator, license, and update history are also unknown.
A subset of a corpus for Urdu text-to-speech synthesis, published on Kaggle. The dataset likely contains audio recordings paired with corresponding text transcripts. Specific details on size, collection method, and contributors are not provided in the available metadata.
A processed subset of an Urdu text-to-speech corpus, published on Kaggle. The dataset likely contains aligned audio recordings and corresponding text transcripts for speech synthesis tasks. Specific details on size, creation date, and original source are not provided in the available metadata.
Mel spectrograms provide a time-frequency representation of audio signals, commonly used for machine learning tasks. This dataset, hosted on Kaggle, likely contains pre-computed mel spectrogram features derived from music audio tracks. The specific source, size, and creation details are not provided in the available metadata.
Azerbaijani Asr Zenfira is a speech dataset hosted on HuggingFace by tahmaz. The dataset card indicates it is intended for automatic speech recognition tasks. Its last update was recorded on February 20, 2026.
Urdu TTS Corpus Subset is a dataset hosted on Kaggle, likely containing audio recordings and corresponding text transcripts for speech synthesis. The dataset's author, size, and specific content details are not provided in the metadata. Users must download the dataset to verify its exact composition and suitability for their projects.
A speech audio dataset combining the LibriSpeech corpus with MUSAN augmentation data. The dataset is published on Kaggle, but specific details on size, creation date, and author are not provided in the metadata. Its content likely contains speech recordings augmented with noise and music samples for machine learning training.
Trained model weights and datasets for the BACHI chord recognition system. The data supports the paper 'BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music' by Mingyang Yao and Ke Chen, accepted for ICASSP 2026. The dataset page was last updated on 2026-01 17.
Synthetic-dysarthric-speech is a dataset containing artificially generated and augmented speech samples simulating dysarthria, a motor speech disorder. It is intended for developing robust automatic speech recognition and semantic understanding systems. The dataset's creator, size, and update date are not specified.
A classification dataset for predicting the commercial success of music tracks on Spotify. The dataset likely contains audio features and metadata to categorize songs into High, Medium, or Low popularity tiers. It was sourced from Kaggle, but details on its creator, size, and specific features are not provided.
A dataset for building music recommendation systems, sourced from the Kaggle platform. The specific content, scale, and features are not detailed in the available metadata. Further details regarding the data's origin, collection method, and temporal coverage are unknown.
ACI-Bench-MedARC evaluates model performance in converting clinical dialogue into structured clinical notes. The dataset includes the benchmark and data from ablation studies testing different transcription methods. It was uploaded by mkieffer to HuggingFace and last updated on 2026-01-18.
More than 16,000 songs from NetEase Cloud Music are included in this multi-label music emotion recognition dataset. MFCC features were extracted from the middle 30 seconds of each song using librosa. The dataset was created by joyfuljune and last updated on 2026-01-29.