Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
lmarena-ai's PPE-GPQA-Best-of-K dataset contains a correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from the GPQA dataset, and the collection is intended for benchmarking and evaluation, not for training. The dataset was last updated on October 22, 2024.
User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.