Paper recorded by Signals 4 on 2026-08-30 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-30 on arXiv · recorded by Signals 4 on 2026-09-01
Category: cs.CV · 计算机视觉 · first seen 2026-09-01
Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized semantic distinctions where top-ranked near misses differ from the true match by a single critical detail. Hard-sample mining can select confusable candidates but cannot construct corrected counterparts; synthetic augmentation can generate novel samples