Signals 4 · free daily AI digest

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

Paper recorded by Signals 4 on 2026-09-01 in cs.CL. Abstract reproduced from arXiv; link to the original below.

Published 2026-09-01 on arXiv · recorded by Signals 4 on 2026-09-02

Category: cs.CL · 自然语言处理 · first seen 2026-09-02

Abstract

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-tra

Read on arXiv →

#143 most recent of 186 cs.CL papers we have recorded · ↑ newer: SDARE-Bench: Evaluating Large Language Models on Conversational Stigma · ↓ older: GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions
Cite this page: Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall: the #143 most recent of 186 cs.CL papers we have recorded (as of 2026-09-01). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/knowledge-distillation-during-mid-training-favors-reasoning-over-factual-recall.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.CL papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions