Paper recorded by Signals 4 on 2026-09-25 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-25 on arXiv · recorded by Signals 4 on 2026-09-28
Category: cs.LG · 机器学习 · first seen 2026-09-28
Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The