Paper recorded by Signals 4 on 2026-08-30 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-30 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.LG · 机器学习 · first seen 2026-09-01
As context lengths scale, attention increasingly becomes a primary computational bottleneck in large language models. Standard Transformers remain powerful but computationally inefficient, as they allocate the same attention budget to every token regardless of its contextual demand. Existing local-global hybrids provide a more efficient alternative by mixing restricted- and full-context attention,