Paper recorded by Signals 4 on 2026-09-14 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-14 on arXiv · recorded by Signals 4 on 2026-09-15
Category: cs.LG · 机器学习 · first seen 2026-09-15
As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals