Sign in to view source links and access this dataset
Description
Approximately 13,500 audio segments totaling around 100 hours of Nepali speech, professionally curated for Automatic Speech Recognition research. The dataset, created by 'tonibirat', contains high-fidelity 16kHz, 16-bit mono WAV files, with segments typically 15-20 seconds long. It was last updated on Hugging Face in January 2026.
Use Cases
Training end-to-end ASR models based on the high-fidelity audio and transcriptions.
Benchmarking model performance on Nepali speech based on the professionally curated test set.
Fine-tuning pre-trained multilingual speech models based on the domain-specific Nepali data.
Studying acoustic features of Nepali speech based on the standardized audio format.
Strengths
Approximately 100 hours of audio provides substantial training material.
~13,500 segments offer a high number of distinct data points.
Audio is standardized to 16kHz, 16-bit, mono WAV format for consistency.
Described as 'professionally curated' and 'high-fidelity', suggesting quality control.
Limitations
Description metadata is limited; actual data quality requires manual inspection after download.
Column-level documentation is absent; field semantics must be inferred after download.
Provenance
Source
huggingface
Collection Method
Professionally curated, method unspecified.
Time Range
null
Freshness
Last updated 2026-01 26 07:13:25; freshness should be verified.
Geography
null
License is unknown; users must verify terms before use.