Paper recorded by Signals 4 on 2026-10-01 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-10-01 on arXiv · recorded by Signals 4 on 2026-10-02
Category: cs.CV · 计算机视觉 · first seen 2026-10-02
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the t