Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
RealMath-Eval is a benchmark for evaluating large language model judges on authentic human mathematical reasoning. The dataset was created by RicharMd and is associated with the paper 'RealMath-Eval: The Evaluation Gap in Judging Human Mathematical Reasoning'. The dataset was last updated on Hugging Face on June 10, 2026.
License information is unknown; users should verify terms of use before downloading.