Paper recorded by Signals 4 on 2026-09-09 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-09 on arXiv · recorded by Signals 4 on 2026-09-10
Category: cs.CV · 计算机视觉 · first seen 2026-09-10
While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true generalization or just domain adaptation. To investigate this gap, we evaluate three AVSR architectures across six conditions: controlled broadcast speech, fixed-grammar utterances, hyper-articulated Lombard speech, read speech from professional lipsp