KL Divergence VS MSE for Knowledge Distillation
Unknown
MSE loss outperforms KL divergence in knowledge distillation by enabling direct logit matching, with sequential distillation and small-tau KL improving noise robustness.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
MSE loss outperforms KL divergence in knowledge distillation by enabling direct logit matching, with sequential distillation and small-tau KL improving noise robustness.
Unknown
A convolution-free vision transformer using a teacher-student strategy with attention-based distillation tokens.
Unknown
LLMLingua2 proposes a task-agnostic prompt compression method using data distillation and a Transformer encoder for token classification.
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, et al.
Gecko is a 1.2B parameter text embedding model that uses two-stage LLM distillation to achieve state-of-the-art retrieval and classification performance.
Unknown
SimCLRv2 presents a semi-supervised learning framework combining unsupervised pretraining, supervised fine-tuning, and distillation with unlabeled data.
Unknown
An open-source family of heterogeneous reasoning models (Nano, Super, Ultra) with dynamic reasoning toggle, trained via NAS, distillation, and RL.
Unknown
Gemma 3 is a multimodal language model with vision understanding, 128k context, and improved performance via distillation and a novel post-training recipe.
Unknown
Pruning and knowledge distillation can compress LLMs 2-4x with up to 40x compute savings and improved performance.
Unknown
Gemma 2 improves open language models by interleaving local-global attention and using knowledge distillation, achieving competitive performance with much larger models.
Unknown
TinyBERT distills BERT into a compact model using attention transfer and two-stage distillation for efficient NLU.
Unknown
Relational knowledge distillation extends traditional distillation by transferring structural relationships between data points from teacher to student.
Xiaohan Xu, Ming Li, Chongyang Tao, et al.
This survey systematically reviews knowledge distillation techniques for transferring capabilities from large proprietary LLMs to smaller models.