Paper recorded by Signals 4 on 2026-09-22 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-22 on arXiv · recorded by Signals 4 on 2026-09-23
Category: cs.CV · 计算机视觉 · first seen 2026-09-23
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-grained mouth representations. A unified joint self-attention mechanism processes all input streams togeth