Paper recorded by Signals 4 on 2026-09-21 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-21 on arXiv · recorded by Signals 4 on 2026-09-22
Category: cs.CV · 计算机视觉 · first seen 2026-09-22
We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically ri