Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference
Paper recorded by Signals 4 on 2026-09-22 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-22 on arXiv · recorded by Signals 4 on 2026-09-23
Category: cs.AI · 人工智能 · first seen 2026-09-23
Abstract
Greedy decoding from large language models is commonly treated as deterministic. We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware. Across our evaluations of six models (1.1B-7B parameters, four families; divergence additionally characterised at 12B) and three benchmarks, 49-100\% of prompts div
Read on arXiv →
Cite this page: Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference: the #19 most recent of 360 cs.AI papers we have recorded (as of 2026-09-22). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/greedy-decoding-is-not-precision-invariant-cross-precision-output-divergence-in-.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free