Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
Paper recorded by Signals 4 on 2026-09-15 in cs.AI. Abstract reproduced from arXiv; link to the original below.
Published 2026-09-15 on arXiv · recorded by Signals 4 on 2026-09-16
Category: cs.AI · 人工智能 · first seen 2026-09-16
Abstract
Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a smaller student. Direct imitation, however, can also transfer the teacher's systematic bias and errors. This challenge is particularly pronounced under covariate shift, when the teacher's reliability on target questions is uncertain and target-domain reward feedback is unavailable. We propose Coupl
Read on arXiv →
Cite this page: Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback: the #51 most recent of 300 cs.AI papers we have recorded (as of 2026-09-15). Source: Signals 4 (Signals API) — https://data.jiangzhang.ca/signals4/t/papers/coupled-calibration-and-learning-mitigating-teacher-bias-in-llm-distillation-wit.html
Free to quote with attribution to “Signals 4 (Signals API)”. Machine-readable:
papers.json
Get 4 AI signals a day by email — free.
Get 4 AI signals a day by email — free