UniSD: Towards a Unified Self-Distillation Framework for Large Language Models

May 7, 2026·
Yiqiao Jin
Yiqiao Jin
,
Yiyang Wang
,
Lucheng Fu
,
Yijia Xiao
,
Yinyi Luo
,
Haoxin Liu
,
B. Aditya Prakash
,
Josiah Hester
,
Jindong Wang
,
Srijan Kumar
· 1 min read
Abstract
UniSD studies how to adapt autoregressive language models through self-distillation without stronger external teachers. It combines multi-teacher agreement, EMA teacher stabilization, token-level contrastive learning, feature matching, and divergence clipping to examine supervision reliability, representation alignment, and training stability. Experiments across six benchmarks and six models identify complementary components for an integrated self-distillation pipeline.
Type
Publication
arXiv preprint arXiv:2605.06597

Overview

UniSD studies how to adapt autoregressive language models through self-distillation without stronger external teachers. It brings together teacher agreement, EMA stabilization, contrastive learning, feature matching, and divergence clipping, examining their individual effects and interactions across six benchmarks and six models.

Yiqiao Jin
Authors
Ph.D. Candidate in Computer Science
My research sits at the intersection of multimodal foundation models, intelligent agents, and spatial/physical intelligence. I build general-purpose models that perceive, reason, learn, and act in interactive virtual and physical environments.