Paper recorded by Signals 4 on 2026-08-30 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-30 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.CV · 计算机视觉 · first seen 2026-09-01
Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response, OSINT verification, and location privacy. However, most efforts to study AI behavior on the task remain limited to static image-based retrieval, classification, and predictions. We argue that faithful recreation of the task should involve embodied n