Sign in to view source links and access this dataset
Description
10,934 real-world audio recordings from Farmer.Chat provide a benchmark for speech-to-text models in agricultural advisory contexts. The dataset is human-annotated and focuses on three Indian languages: Hindi, Telugu, and Odia. Bullseye-4 created this resource, which was last updated on March 20, 2026.
Use Cases
Benchmarking ASR model performance based on real-world agricultural speech audio
Training domain-specific speech recognition models based on human-annotated transcriptions
Evaluating multilingual ASR systems based on the Hindi, Telugu, and Odia language focus
Studying speech patterns in agricultural advisory scenarios based on the Farmer.Chat source
Strengths
10,934 audio recordings provide a substantial corpus for benchmarking
Human-annotated transcriptions likely ensure high-quality ground truth labels
Focus on three specific languages (Hindi, Telugu, Odia) offers targeted multilingual data
Limitations
Column-level documentation is absent; field semantics must be inferred after download
Row count is unknown, which may limit suitability assessment
Data may reflect geographic bias inherent to the Farmer.Chat source and selected languages
Provenance
Source
Farmer.Chat
Collection Method
Likely collected from real-world agricultural advisory interactions.
Freshness
Last updated 2026-03-20 12:50:45; freshness should be verified
Geography
Likely India, given the focus on Hindi, Telugu, and Odia languages.
License is unknown; terms of use must be verified before application.