Paper recorded by Signals 4 on 2026-09-15 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-15 on arXiv · recorded by Signals 4 on 2026-09-16
Category: cs.CV · 计算机视觉 · first seen 2026-09-16
Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We study this design choice directly. A central finding is that changing the 2D backbo