Sign in to view source links and access this dataset
Description
Arabic instruction-tuning data combining instruction-following pairs, instruction descriptions, freeform responses, and quality control data. The dataset contains over 3,300 records designed for fine-tuning Arabic language models, created by ArSyra and last updated in March 2026.
Use Cases
Supervised fine-tuning (SFT) of Arabic LLMs based on the described instruction-following pairs.
Training models to handle Arabic dialectal variations based on the dialectal instruction pairs mentioned.
Improving instruction-following capabilities in Arabic LLMs based on the described instruction descriptions and responses.
Benchmarking or evaluating the quality of Arabic LLM outputs using the described quality control data.
Strengths
Over 3,300 records specifically for Arabic LLM fine-tuning.
Includes multiple data components: instruction-following pairs, descriptions, freeform responses, and quality control data.
Focuses on Arabic dialects, a noted gap in NLP resources.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known (3,300+), but the exact number and data scale are not fully detailed.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
ArSyra
Freshness
Last updated 2026-03-15 00:00:35; freshness should be verified.
Geography
Arabic-speaking regions, with a focus on dialectal data.
License is unknown; users must verify terms of use before downloading.