研究提出 Switch Distillation:中训练阶段知识蒸馏更利于推理而非事实记忆

内容摘要
研究指出,基于Logit的知识蒸馏(KD)在训练较小语言模型时,通过从更强的教师模型中获取监督,但其效益是否在训练阶段保持一致尚不明确。研究发现,在训练中期,即自监督学习过程中,使用经过后训练的教师模型进行的前向KL蒸馏(标准KD公式)表现出了与预训练阶段不同的行为。尽管前向KD在预训练阶段相对于标准下一个标记预测(NTP)同时提高了推理和事实回忆能力,但在训练中期,尽管推理能力持续提升,事实回忆能力却有所下降。这种阶段依赖性源于教师在不同数据领域上的信心不对称以及学生知识状态的演变:教师在程序性数据上的信心更高,而学生在训练早期就获得了低熵的事实知识。为了缓解这种不平衡,研究者提出了Switch Distillation,这是一种简单的训练中期目标,它通过教师预测熵作为轻量级路由信号,在教师有信心的地方进行蒸馏,否则则回退到交叉熵。Switch Distillation在教师规模上始终优于现有的蒸馏目标,相对于标准NTP,它在推理性能上提高了1.61-1.71倍,在知识和常识性能上提高了1.13-1.19倍,同时保留了96.7-96.8%的事实回忆。重要的是,这些效益在训练后仍然存在:Switch Distillation缩小了事实回忆差距,同时保持了推理和知识及常识性能的1.25-1.32倍和1.13-1.20倍的提升。
研究指出,基于Logit的知识蒸馏(KD)在训练较小语言模型时,通过从更强的教师模型中获取监督,但其效益是否在训练阶段保持一致尚不明确。研究发现,在训练中期,即自监督学习过程中,使用经过后训练的教师模型进行的前向KL蒸馏(标准KD公式)表现出了与预训练阶段不同的行为。尽管前向KD在预训练阶段相对于标准下一个标记预测(NTP)同时提高了推理和事实回忆能力,但在训练中期,尽管推理能力持续提升,事实回忆能力却有所下降。这种阶段依赖性源于教师在不同数据领域上的信心不对称以及学生知识状态的演变:教师在程序性数据上的信心更高,而学生在训练早期就获得了低熵的事实知识。为了缓解这种不平衡,研究者提出了Switch Distillation,这是一种简单的训练中期目标,它通过教师预测熵作为轻量级路由信号,在教师有信心的地方进行蒸馏,否则则回退到交叉熵。Switch Distillation在教师规模上始终优于现有的蒸馏目标,相对于标准NTP,它在推理性能上提高了1.61-1.71倍,在知识和常识性能上提高了1.13-1.19倍,同时保留了96.7-96.8%的事实回忆。重要的是,这些效益在训练后仍然存在:Switch Distillation缩小了事实回忆差距,同时保持了推理和知识及常识性能的1.25-1.32倍和1.13-1.20倍的提升。

Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值