Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Spider DPO 1040 is a compact training dataset containing 1,040 preference pairs for Direct Preference Optimization, derived from frontier-model disagreements on the Spider V1 benchmark. It also includes 7,000 supervised training examples from Spider formatted for use with LLaMA-Factory. The dataset was created by jk200201 and was last updated on July 5, 2026.
The full dataset description is hosted externally on Hugging Face. The dataset is designed for use with a specific companion LoRA adapter.