Sign in to view source links and access this dataset
Description
Moroccan Darija ASR Dataset Split is a speech corpus for Automatic Speech Recognition, published on the Hugging Face platform by mohamedmou. The dataset was last updated on May 1, 2026, but its specific size, content, and collection methodology are not detailed in the available metadata.
Use Cases
Train an Automatic Speech Recognition (ASR) model for Moroccan Darija (inferred from domain, verify after download)
Benchmark ASR performance on low-resource language varieties (inferred from domain, verify after download)
Fine-tune pre-trained multilingual speech models for a specific dialect (inferred from domain, verify after download)
Strengths
Published on the Hugging Face platform, facilitating access for the ML community.
Authored by a named contributor (mohamedmou), providing a point of contact.
Limitations
Metadata is minimal; actual content requires verification after download.
Row count, audio duration, and column-level documentation are unknown.
The dataset's license, collection method, and geographic/speaker diversity are unspecified.
Provenance
Source
huggingface
Freshness
Last updated 2026-05-01 21:36:53
Geography
Morocco (inferred from title)
License is unknown; users must verify terms before commercial use.