Paper recorded by Signals 4 on 2026-09-29 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-29 on arXiv · recorded by Signals 4 on 2026-09-30
Category: cs.AI · 人工智能 · first seen 2026-09-30
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological visi