Paper recorded by Signals 4 on 2026-09-08 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-08 on arXiv · recorded by Signals 4 on 2026-09-09
Category: cs.CL · 自然语言处理 · first seen 2026-09-09
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These evaluations do not fully capture how visual tokens behave when modeled jointly with text. We build a controlled pure-autoregressive testbed and track task-specific validation losses during multimodal continual pretraini