Paper recorded by Signals 4 on 2026-09-16 in cs.CL. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-16 on arXiv · recorded by Signals 4 on 2026-09-17
Category: cs.CL · 自然语言处理 · first seen 2026-09-17
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or i