Paper recorded by Signals 4 on 2026-09-23 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-23 on arXiv · recorded by Signals 4 on 2026-09-24
Category: cs.CL · 自然语言处理 · first seen 2026-09-24
Diffusion Language Models (DLMs) have attracted significant attention for their strong reasoning ability. However, under a bidirectional attention mechanism, DLMs operate over an exponentially large exploration space compared to autoregressive models (ARMs), making it challenging to focus on reasoning-guiding tokens under random masking. We define causal shortcuts as token chains that cover the fu