Paper recorded by Signals 4 on 2026-09-02 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-02 on arXiv · recorded by Signals 4 on 2026-09-03
Category: cs.CV · 计算机视觉 · first seen 2026-09-03
We introduce SolarWM, a fully open foundation for building interactive video world models from data preparation through long-horizon inference. Training across heterogeneous data sources and video backbones is challenging: datasets differ in temporal scale, camera geometry, visual quality, motion, and captioning styles, while video generators use distinct representations and architectures. Naive d