Paper recorded by Signals 4 on 2026-09-14 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-14 on arXiv · recorded by Signals 4 on 2026-09-15
Category: cs.AI · 人工智能 · first seen 2026-09-15
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan