Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
4.8 million collective human preferences compare the helpfulness of two responses to questions or instructions. The dataset spans 129 diverse subject areas, from cooking to legal advice, and is an extended version of the original 385K SHP dataset. Created by stanfordnlp and updated in January 2024, it is intended for training RLHF reward models and NLG evaluation models.
License is unknown; terms of use must be verified before application.