Signals 4 · free daily AI digest

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Paper recorded by Signals 4 on 2026-09-03 in cs.CV. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-03 on arXiv · recorded by Signals 4 on 2026-09-04

Category: cs.CV · 计算机视觉 · first seen 2026-09-04

Abstract

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), togeth

Read on arXiv →

#146 most recent of 237 cs.CV papers we have recorded · ↑ newer: Principia: Relational Physics Tests for Video Models · ↓ older: Zero-Shot Novel Depth Synthesis Using 3D Foundation Models Scene Repre
Cite this page: Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States: the #146 most recent of 237 cs.CV papers we have recorded (as of 2026-09-03). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/puffin-world-scaling-a-unified-multimodal-model-with-native-3d-world-states.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CV papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions