Signals 4 · free daily AI digest

Minimally Invasive Steering of Language Models

Paper recorded by Signals 4 on 2026-09-24 in cs.AI. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-24 on arXiv · recorded by Signals 4 on 2026-09-25

Category: cs.AI · 人工智能 · first seen 2026-09-25

Abstract

Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The res

Read on arXiv →

#10 most recent of 400 cs.AI papers we have recorded · ↑ newer: Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parame · ↓ older: Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
Cite this page: Minimally Invasive Steering of Language Models: the #10 most recent of 400 cs.AI papers we have recorded (as of 2026-09-24). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/minimally-invasive-steering-of-language-models.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.AI papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions