Paper recorded by Signals 4 on 2026-09-30 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-30 on arXiv · recorded by Signals 4 on 2026-10-01
Category: cs.LG · 机器学习 · first seen 2026-10-01
Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object recognition, common-sense reasoning, and contextual understanding, capabilities that align with the nuanced requirements of social robot navigation. However, it remains unclear wheth