arXiv AI

On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures

arXiv:2607. 15919v1 Announce Type: cross Abstract: Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original training data.

arXiv AI
Aug 11

UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation

arXiv:2608. 09287v1 Announce Type: cross Abstract: Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset.

By Xuewan He, Tong Chu, Zihan Cheng, Yuchen Su, Qianxin Xia, Guoming Lu, Jielei Wang, Wen Li
arXiv AI
Aug 26

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.

By Irene Trigueros-Lorca, Leonardo Concepci\'on, Christian Wagner, Isaac Triguero, Daniel Molina
arXiv Computer Vision
2d ago

Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring

Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.

By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
arXiv AI
Jul 14

Stable On-Policy Distillation through Adaptive Target Reformulation

arXiv:2601. 07155v3 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models; however, conventional supervised KD often suffers from a distribution mismatch between training and inference.

By Ijun Jang, Jewon Yeom, Juan Yeo, Hyunggyu Lim, Taesup Kim
arXiv Computer Vision
Aug 27

CloSeR: Unified Relational Distillation from Closed-Set Teachers for Category Discovery

CloSeR is a plug‑and‑play framework that enhances Generalized Category Discovery (GCD) by injecting closed‑set relational knowledge from a lightweight teacher model. The teacher is built by fine‑tuning adapters on labeled data while keeping the backbone frozen, preserving pretrained priors. Unified Relational Distillation then transfers both global sample‑to‑prototype and local sample‑to‑sample relations to the GCD task, reducing optimization interference and improving performance across six benchmarks with DINO and DINOv2 backbones.

By Yuanpei Liu, Zhenqi He, Jialu Tang, Kai Han
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
arXiv Machine Learning
4d ago

Data Unlearning via Inverse Distillation

arXiv:2609.36099v1 Announce Type: new Abstract: Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce u...

By Aleksei Leonov, Nikita Kornilov, Zhenhe Zhang, Evgeny Burnaev, Iaroslav Koshelev, Alexander Korotin