Paper recorded by Signals 4 on 2026-09-18 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-18 on arXiv · recorded by Signals 4 on 2026-09-21
Category: cs.CV · 计算机视觉 · first seen 2026-09-21
Current Vision-Language-Action (VLA) models for autonomous driving operate primarily through feedforward inference across the perception--reasoning--planning hierarchy. While modern architectures maintain temporal recurrence within the perceptual module, early perception remains blind to downstream reasoning and navigation goals, processing visual inputs agnostically without prioritizing cues info