Paper recorded by Signals 4 on 2026-09-22 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-22 on arXiv · recorded by Signals 4 on 2026-09-23
Category: cs.CV · 计算机视觉 · first seen 2026-09-23
Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challenge. We argue that the root cause lies in the entanglement of the Encoder--Decoder and Codebook train