Preprint2026
Reward under attack: Analyzing the robustness and hackability of process reward models
Unknown
This paper demonstrates that state-of-the-art Process Reward Models (PRMs) are systematically exploitable, revealing critical robustness vulnerabilities in LLM reasoning pipelines.
0Mar 1, 2026Reasoning
arXiv