Paper recorded by Signals 4 on 2026-09-30 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-30 on arXiv · recorded by Signals 4 on 2026-10-01
Category: cs.AI · 人工智能 · first seen 2026-10-01
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{pe