Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
AllenAI created this dataset of 19,890 synthetically generated preference examples to enhance models' precise instruction-following capabilities. It contains chosen and rejected response pairs, intended for preference tuning methods like PPO and DPO. The dataset was last updated on November 21, 2024.
License is unknown; terms of use must be verified before application.