Paper recorded by Signals 4 on 2026-09-23 in cs.CV. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-23 on arXiv · recorded by Signals 4 on 2026-09-24
Category: cs.CV · 计算机视觉 · first seen 2026-09-24
Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. However, quantization affects different layer types in inconsistent ways, so identifying where accuracy loss is minimized and latency reduction is maximized is critical, as the effect accumulates over a ful