Paper recorded by Signals 4 on 2026-09-15 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-15 on arXiv · recorded by Signals 4 on 2026-09-16
Category: cs.LG · 机器学习 · first seen 2026-09-16
Large language models (LLMs) trained on next-token prediction exhibit remarkable in-context learning (ICL) abilities, yet the representations that support ICL remain poorly understood. We consider such representations in a controlled setting: prompting LLMs with data emitted from hidden Markov models (HMMs) and probing for the corresponding belief state -- the posterior distribution over the HMM's