Sign in to view source links and access this dataset
Description
9,304 spoken Arabic prompts from 93 users interacting with an ASR and LLM-based assistant. The WASIL dataset, created by QCRI, captures in-the-wild interactions across multiple dialects and countries, including explicit user feedback signals like likes, dislikes, and scalar scores. The dataset was last updated on Hugging Face in May 2026.
Use Cases
Training dialect-aware Arabic speech recognition models based on in-the-wild spoken prompts.
Evaluating LLM assistant responses for Arabic users based on explicit like/dislike feedback.
Analyzing user satisfaction and interaction patterns across different Arabic dialects.
Fine-tuning conversational AI systems using real-world, annotated Arabic dialogue turns.
Strengths
Contains 9,304 dialogue turns, providing a substantial corpus of spoken interactions.
Includes explicit user feedback signals (like/dislike and scalar scores) for 93 users.
Spans multiple Arabic dialects and countries, suggesting linguistic diversity.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known but other scale details like file size and formats are unknown.
Data may reflect geographic and demographic bias inherent to its collection platform and user base.
Provenance
Source
QCRI
Collection Method
Collected from in-the-wild interactions with an ASR → LLM assistant.
Time Range
null
Freshness
Last updated 2026-05-19 16:50:32; freshness should be verified.
Geography
Spans multiple countries, but specific locations are not detailed.
License information is unknown; terms of use must be verified before application.