Signals 4 · free daily AI digest

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Paper recorded by Signals 4 on 2026-09-29 in cs.AI. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-29 on arXiv · recorded by Signals 4 on 2026-09-30

Category: cs.AI · 人工智能 · first seen 2026-09-30

Abstract

Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natu

Read on arXiv →

#3 most recent of 460 cs.AI papers we have recorded · ↑ newer: STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State · ↓ older: Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity
Cite this page: LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization: the #3 most recent of 460 cs.AI papers we have recorded (as of 2026-09-29). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/leapquant-efficient-linear-attention-with-accurate-recurrent-state-quantization.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.AI papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions