Sign in to view source links and access this dataset
Description
Neyshekar is an open Persian speech dataset containing 40,008 transcribed audio clips totaling approximately 63 hours. The data was collected from native Persian speakers through a community-driven crowdsourcing platform. This release, V4.1, was created by shekar-ai and last updated on June 15, 2026.
Use Cases
Train Persian automatic speech recognition models based on 40,008 transcribed audio clips.
Develop text-to-speech systems for Persian using native speaker audio.
Conduct speech representation learning research with a dedicated Persian corpus.
Benchmark speech processing algorithms on a crowdsourced Persian dataset.
Strengths
Contains 40,008 transcribed audio samples, providing a substantial corpus.
Totals approximately 63 hours of speech data for model training.
Audio is recorded in a consistent format: mono, 16-bit PCM, sampled at 16 kHz.
Data is sourced from native Persian speakers, supporting language-specific research.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
shekar-ai via Hugging Face
Collection Method
Collected from native Persian speakers through a community-driven crowdsourcing platform.
Freshness
Last updated 2026-06-15 10:39:37; freshness should be verified.
Geography
Persian language speakers; specific geographic coverage is not detailed.
License is unknown; terms of use must be verified before application.