UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
May 7, 2026·
,,,,,,,,,·
1 min read
Yiqiao Jin
Yiyang Wang
Lucheng Fu
Yijia Xiao
Yinyi Luo
Haoxin Liu
B. Aditya Prakash
Josiah Hester
Jindong Wang
Srijan Kumar

Abstract
UniSD studies how to adapt autoregressive language models through self-distillation without stronger external teachers. It combines multi-teacher agreement, EMA teacher stabilization, token-level contrastive learning, feature matching, and divergence clipping to examine supervision reliability, representation alignment, and training stability. Experiments across six benchmarks and six models identify complementary components for an integrated self-distillation pipeline.
Type
Publication
arXiv preprint arXiv:2605.06597
Overview
UniSD studies how to adapt autoregressive language models through self-distillation without stronger external teachers. It brings together teacher agreement, EMA stabilization, contrastive learning, feature matching, and divergence clipping, examining their individual effects and interactions across six benchmarks and six models.