Paper recorded by Signals 4 on 2026-09-25 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-25 on arXiv · recorded by Signals 4 on 2026-09-28
Category: cs.LG · 机器学习 · first seen 2026-09-28
Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic O