Signals 4 · free daily AI digest

On-Demand Attention: Language Models Know When to Recall

Paper recorded by Signals 4 on 2026-09-17 in cs.CL. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-17 on arXiv · recorded by Signals 4 on 2026-09-18

Category: cs.CL · 自然语言处理 · first seen 2026-09-18

Abstract

Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA),

Read on arXiv →

#3 most recent of 186 cs.CL papers we have recorded · ↑ newer: JEPA-Anything: Learning Predictive Models across Different Worlds · ↓ older: Summarization Bias: The Directional Collapse of Objective Projection i
Cite this page: On-Demand Attention: Language Models Know When to Recall: the #3 most recent of 186 cs.CL papers we have recorded (as of 2026-09-17). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/on-demand-attention-language-models-know-when-to-recall.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CL papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions