Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
GM-PRM-20K is a training dataset for a Generative Multimodal Process Reward Model, as described in arXiv:2508.04088. It contains full multi-step solutions to multimodal math problems, used to train the model zijinghuafen/GM-PRM. The dataset was accepted at the 4th Workshop on Advances in Language and Vision Research (ALVR), in conjunction with ACL 2026.
License is unknown, which may restrict usage; users should verify licensing terms before application.