Paper recorded by Signals 4 on 2026-09-21 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-21 on arXiv · recorded by Signals 4 on 2026-09-22
Category: cs.AI · 人工智能 · first seen 2026-09-22
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's