Paper recorded by Signals 4 on 2026-09-02 in cs.LG. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-02 on arXiv · recorded by Signals 4 on 2026-09-03
Category: cs.LG · 机器学习 · first seen 2026-09-03
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications. We instead pair E2M1 payloads with unsigned E5M3 (\ue{}) blo