Paper recorded by Signals 4 on 2026-09-17 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-17 on arXiv · recorded by Signals 4 on 2026-09-18
Category: cs.LG · 机器学习 · first seen 2026-09-18
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present