Paper recorded by Signals 4 on 2026-09-29 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-29 on arXiv · recorded by Signals 4 on 2026-09-30
Category: cs.LG · 机器学习 · first seen 2026-09-30
Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However,