Paper recorded by Signals 4 on 2026-09-15 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-15 on arXiv · recorded by Signals 4 on 2026-09-16
Category: cs.CV · 计算机视觉 · first seen 2026-09-16
Vision-language models (VLMs) achieve strong visual question answering (VQA) performance, but processing large cluttered images is computationally expensive when only a small region is relevant. Electroencephalography (EEG) signals, which capture human neural responses to visual stimuli, can provide a human-derived semantic cue about the region of interest (ROI). However, EEG-guided visual categor