Hugging Face Trending Papers

Grid-Preserving Knowledge Distillation: Transferring Convolutional Inductive Bias to Vision Transformers under Data Scarcity

Vision Transformers underperform convolutional networks when training data is scarce, and distilling convolutional inductive biases from a CNN teacher is an effective remedy that leaves the deployed model unchanged. General-purpose feature distillation, however, transfers little in this setting.

arXiv AI
Aug 26

Too much of a good thing -- when knowledge distillation promotes overfitting, and how to avoid it

The paper investigates how knowledge distillation (KD) applied at intermediate layers of a neural network can affect overfitting and model performance. While traditional KD focuses on the final output, this study explores block‑wise KD across eleven datasets, finding that on standard datasets the last block suffices, but on fine‑grained, data‑scarce settings intermediate supervision significantly improves accuracy. The authors also analyze optimal supervision granularity using attention maps, Centered Kernel Alignment, and Grad‑CAM, and examine teacher‑student fine‑tuning strategies.

By Irene Trigueros-Lorca, Leonardo Concepci\'on, Christian Wagner, Isaac Triguero, Daniel Molina
arXiv Computer Vision
Sep 25

IronViT: Toward Efficient Generalist Visual Representation Learning

IronViT proposes a new approach to building efficient generalist vision encoders by first consolidating the knowledge of multiple specialist teachers into a softmax attention bridge and then transferring this consolidated representation to a hybrid softmax‑linear attention architecture. This two‑stage distillation process, supported by a curated data pipeline, allows the model to capture semantic, spatial, language‑aligned, and action‑relevant cues while avoiding the high‑resolution cost of traditional softmax attention. Across tasks such as recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT matches or exceeds the performance of leading specialist and generalist encoders, with the hybrid encoder offering increasing efficiency at higher resolutions.

By Jiaxi Huang, Yueqi Hu, Xin Zhu, Xiaopeng Zhang, Huiting Qiao, Yanglin Zhang, Zefeng Ji, Rongxue Li, Yifei Xu, Huiying Yu, Wei Liu, Jiayin Zheng, Yinggan Xu, Peipeng Chen, Yin Zhang, Jian Yao
arXiv Computation and Language
Sep 10

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

arXiv:2609.10321v1 Announce Type: new Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in curren...

By Hongyuan Zhang, Xianda Guo, Yanlun Peng, Qianlong Yang, Yubin Guo, Pinhan Fu, Mulin Chen, Xiaozhen Qiao, Ping Luo
arXiv Computer Vision
Aug 27

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

MLLMCLIP introduces a heterogeneous distillation framework that transfers multimodal knowledge from a generative Multimodal Large Language Model (MLLM) teacher directly into a discriminative CLIP student, eliminating the need for synthetic hard negatives. The method uses an attention-based per-layer token selection and a CKA-based distillation loss to bridge architectural differences between the two models. As a result, MLLMCLIP achieves state‑of‑the‑art compositional accuracy and improves zero‑shot classification and image‑text retrieval performance.

By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
arXiv Computer Vision
Sep 3

Breaking the Geometric Bottleneck: Contrastive Expansion in Asymmetric Cross-Modal Distillation

The paper investigates how knowledge distillation from Vision Transformers to smaller CNNs can cause dimensional collapse in the student’s representation space. Using SVD and Shannon entropy, the authors show that cosine‑based distillation leads to a drastic reduction in effective rank, while adding an InfoNCE objective can double the rank but harms downstream accuracy due to signal dilution. They further demonstrate that a label‑aware contrastive objective (Supervised Contrastive distillation) can maintain or improve accuracy without unnecessary rank expansion, indicating that effective rank alone is not a reliable indicator of representation quality.

By Kabir Thayani