Paper recorded by Signals 4 on 2026-10-01 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-10-01 on arXiv · recorded by Signals 4 on 2026-10-02
Category: cs.AI · 人工智能 · first seen 2026-10-02
Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT