Self-distillation—where a single LLM acts as both teacher and student—is a practical way to train reasoning models, but it risks "privileged information leakage," where the student learns shortcuts based on information it won't have at test time. DemoPSD addresses this by steering the student toward a balanced target between its own distribution and the teacher's, dynamically adjusting the blend based on how much they disagree at each token. This preserves exploration while reducing leakage. Tested on scientific reasoning benchmarks, DemoPSD outperforms existing methods like GRPO, offering a more robust training recipe for building capable reasoning models in specialized domains.
Authors: Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song
Paper: https://arxiv.org/abs/2607.02502v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
