Sign in to view source links and access this dataset
Description
A Mongolian-language subset of the Mozilla Common Voice speech recognition dataset, containing 6,018 audio samples totaling 9.12 hours. The data is split into training, validation, and test sets, with average clip durations between 5.14 and 5.73 seconds. It was uploaded by user 'bilguun' to Hugging Face and last updated on April 13, 2026.
Use Cases
Train Mongolian automatic speech recognition (ASR) models based on the described audio samples.
Fine-tune multilingual speech models using the provided Mongolian audio data.
Benchmark speech recognition performance on the described validation and test splits.
Study phonetic or acoustic properties of the Mongolian language using the audio samples.
Strengths
Contains 6,018 audio samples, providing a foundation for model training.
Offers a structured split with 2,188 training, 1,896 validation, and 1,934 test samples for evaluation.
Total duration of 9.12 hours provides a quantifiable volume of speech data.
Average clip durations are provided per split, suggesting consistent sample lengths.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Last updated 2026-04-13 08:01:09; freshness should be verified.
Provenance
Source
Mozilla Common Voice project.
Collection Method
Likely crowd-sourced audio contributions.
Freshness
Last updated 2026-04-13 08:01:09.
Geography
Mongolia or Mongolian-speaking regions.
License is unknown; terms of use must be verified before application.