Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
ArxivMathGradingBench is a benchmark of 35 arXiv mathematics papers, each containing an error later corrected by the authors. The dataset provides paper metadata and annotations for the location of the errors. It was introduced by LukeBailey181Pub to evaluate large language model proof verifiers on research-level mathematics.
The dataset provides metadata and annotations; downloading the paper PDFs requires following separate instructions.