Signals 4 · free daily AI digest

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Paper recorded by Signals 4 on 2026-10-01 in cs.LG. Abstract reproduced from arXiv; link to the original below.

Published 2026-10-01 on arXiv · recorded by Signals 4 on 2026-10-02

Category: cs.LG · 机器学习 · first seen 2026-10-02

Abstract

Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostic

Read on arXiv →

#8 most recent of 362 cs.LG papers we have recorded · ↑ newer: Decoding Looped Transformers Better for (Almost) Free · ↓ older: Effective Resistance and Graph Neural Network Reliability in Tissue-Sp
Cite this page: From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation: the #8 most recent of 362 cs.LG papers we have recorded (as of 2026-10-01). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/from-gradients-to-capabilities-understanding-multi-teacher-on-policy-distillatio.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable: papers.json
Related: More cs.LG papers · arXiv signals · All papers · Today in AI
Get 4 AI signals a day by email — free.
Subscribe free → See all plans →
Get 4 AI signals a day by email — free
All models · All repos · By company · Daily editions