
SARA is a unified RAG framework that balances local factual precision with global coverage by combining natural-language spans with compact semantic compression vectors, achieving consistent gains under strict context budgets.
Jul 1, 2026

A unified study of LLM self-distillation that combines teacher agreement, EMA stabilization, contrastive learning, feature matching, and divergence clipping to improve adaptation without stronger external teachers.
May 30, 2026

A unified study of LLM self-distillation that combines teacher agreement, EMA stabilization, contrastive learning, feature matching, and divergence clipping to improve adaptation without stronger external teachers.
May 7, 2026

Selective and Adaptive Retrieval-augmented Generation with Context Compression. A unified RAG framework that combines fine-grained natural-language spans with compact semantic compression vectors under strict context budgets. Accepted at ACL 2026 main conference.
Mar 15, 2026

Distills multi-agent debate into one LLM through reasoning-enhanced fine-tuning, trajectory-based augmentation, and process-aware distillation, moving computation from inference to training.
Feb 3, 2026
An efficient approach to probing LLM knowledge that adapts pre-trained embeddings to query model knowledge with substantially reduced compute.
Aug 8, 2025

A survey of efficient LLM training organized around data-centric techniques — selection, mixing, ordering, and synthesis — and their trade-offs with compute and downstream performance.
Jul 31, 2025