Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
348 under-served languages are represented in this collection of spontaneous speech recordings and transcriptions. The corpus was collected by Meta FAIR's Omnilingual ASR project for training automatic speech recognition and spoken language identification models. It was last updated on the Hugging Face platform in December 2025.
License is unknown; users should verify terms of use before downloading.