Paper recorded by Signals 4 on 2026-08-28 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-08-28 on arXiv · recorded by Signals 4 on 2026-08-31
Category: cs.AI · 人工智能 · first seen 2026-08-31
This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias,