Paper recorded by Signals 4 on 2026-09-28 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-28 on arXiv · recorded by Signals 4 on 2026-09-29
Category: cs.AI · 人工智能 · first seen 2026-09-29
In reinforcement learning with verifiable rewards (RLVR), imperfect verifiers can reward incorrect responses, creating opportunities for reward hacking. Using gradient flow with a fixed verifier, we characterize the conditions under which reward rises while correctness falls. We then show that the observations available during RLVR are, in general, insufficient to detect or identify accepted error