Paper recorded by Signals 4 on 2026-09-11 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-11 on arXiv · recorded by Signals 4 on 2026-09-14
Category: cs.CV · 计算机视觉 · first seen 2026-09-14
Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, treating videos as image sequences rather than observations of an underlying dyna