Loading...
Loading...
Speech recognition, text-to-speech, speaker identification, music classification, audio event detection
2,587 datasets
Audio recordings collected in community settings in Senegal cover topics including family planning, healthcare access, pregnancy practices, and cultural beliefs around maternal and reproductive health. The dataset was created by YUXCulturalAILab and last updated on March 19,我们发现了一个问题。 2026. Recordings were captured using mobile devices or portable recorders in natural conversational conditions, and all transcriptions were manually verified.
A Yoruba language speech corpus likely intended for text-to-speech applications. The dataset is hosted on Kaggle and appears to be associated with a Jupyter notebook. The number of speakers, recording hours, and specific collection details are unknown.
A bilingual text-to-speech dataset containing Hebrew and English audio generated by male and female speakers. Audio files have been resampled to 44.1kHz and time-stretched to a slower speed. The dataset was created by author notmax123 and last updated on March 30, 2026.
Asru Data is a dataset uploaded to HuggingFace by author closerG. The dataset was last updated on 2026-05-14. Its specific content and scale are not detailed in the provided metadata.
ViMedCSS is a dataset for medical speech recognition, sourced from the HuggingFace platform and hosted on Kaggle. The dataset's specific size, format, and collection details are not provided in the available metadata. Its primary application appears to be in training or evaluating automated speech recognition systems for clinical or healthcare settings.
MOSS TTSD SGLang assets are hosted on Kaggle. The dataset likely contains audio and text assets for text-to-speech synthesis. Specific details on size, format, and creation are unavailable from the provided metadata.
Regulatory text covers structural integrity, fire safety, and energy conservation for all new construction, renovation, and demolition projects in Massachusetts. The code is written by the State Board of Regulations and Standards and administered locally by certified building inspectors. The dataset originates from the SCIOPS organization via the NASA Earthdata platform.
Irodori TTS Voice Clones is a collection of 2.99 million voice clones for text-to-speech synthesis. It was created by SynDataLab and references the SynDataLab/irodori-refs-10k dataset for source audio. The dataset was last updated on April 23, 2026.
Fall 2003 documentation details the Massachusetts air quality program's implementation of federal and state Clean Air Acts. The dataset includes regulatory procedures, application forms, fee structures, and review timelines for construction permits. It was published by SCIOPS in 2003.
Environmental Protection Agency's BEACH Program data focuses on improving public health for beachgoers through five key areas, including pollution prediction and faster water testing. The program, sponsored by the EPA and managed by SCIOPS, provides information on coastal water quality. Specific contact information is available for data related to Massachusetts beaches.
10,934 real-world audio recordings from Farmer.Chat provide a benchmark for speech-to-text models in agricultural advisory contexts. The dataset is human-annotated and focuses on three Indian languages: Hindi, Telugu, and Odia. Bullseye-4 created this resource, which was last updated on March 20, 2026.
Google Search Console normalized data from the Tably.es marketplace for May 2026. The dataset likely contains aggregated search performance metrics for the platform. The author, organization, and specific data volume are unknown.
Vegetation field plots at Scotts Bluff National Monument were visited, described, and documented in a digital database. The database consists of three parts: Physical Descriptive Data, Species Listings, and Strata Descriptive Data. Information for this metadata was obtained from a USGS site and put into NASA Directory Interchange Format.
DBp is a multimodal dataset from the Music-in-Medicine program, recording a Dueling Brains performance. The data includes audio, video, and tabular file formats, totaling approximately 9.8 GB in size. It is openly licensed under CC-BY-4.0 and authored by Maxine Annel Pacheco-Ramírez.
MHp provides a 5.8 GB multimodal dataset capturing a live 'Musical Healing' performance from the Music-in-Medicine program. It likely contains synchronized electroencephalogram (EEG) brain activity recordings and audio data, such as piano music. This dataset supports research into the neurological and physiological effects of therapeutic music interventions.
Block-scale rooftop solar technical potential estimates for the city of Orlando, Florida, derived from LiDAR and national parcel data. It includes developable roof area and technical potential in kilowatts, along with the most common building use and occupancy type per block.
21,421 cleaned Georgian speech samples totaling 35 hours were curated by NMikka from Mozilla Common Voice 19.0 in 2026. The collection features 24 kHz mono WAV audio from 12 speakers specifically filtered for speech synthesis and recognition tasks.
Contains audio clips for training a model to recognize the keyword 'Sam'. Each clip is labeled as positive (contains 'Sam') or negative (phonetically similar words). The dataset includes varied speaking styles, speeds, and intonations.
Maxine Annel Pacheco-Ramírez's dataset contains multimodal recordings from a Music-in-Medicine program performance titled 'A Musical Dialogue'. The data includes brain activity, audio, and video, but emotional ratings for participant 5 are missing. It is a large dataset, approximately 7.66 GB in size, and is available under a CC-BY-4.0 license.
A parallel speech corpus containing audio recordings paired with text transcripts for the Gojjam dialect of Amharic. It is curated by leyu-amharic to support speech technology research. The dataset was last updated in March 2026.