Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
Paper recorded by Signals 4 on 2026-09-16 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-16 on arXiv · recorded by Signals 4 on 2026-09-17
Category: cs.LG · 机器学习 · first seen 2026-09-17
Abstract
As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to understand and discover the range of hacking behaviors a model displays. In particular, we find that
Read on arXiv →
Cite this page: Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations: the #17 most recent of 215 cs.LG papers we have recorded (as of 2026-09-16). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/monitoring-and-discovering-reward-hacking-with-internal-representations-during-l.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free