Paper recorded by Signals 4 on 2026-09-02 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-02 on arXiv · recorded by Signals 4 on 2026-09-03
Category: cs.LG · 机器学习 · first seen 2026-09-03
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed la