arXiv AI

Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

arXiv Computer Vision
Sep 11

CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

CLIP-RD introduces a relational distillation framework for efficient CLIP knowledge distillation, featuring Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD aligns intra‑modal similarity distributions between teacher and student, while XRD aligns cross‑modal similarity distributions to enforce bidirectional symmetry. This joint modeling of multidirectional relational structures improves the student’s embedding geometry, yielding a 1.8%p performance gain over CLIP‑KD across various architectures, tasks, and corruption settings with minimal training‑time overhead.

By Jeannie Chung, Hanna Jang, Ingyeong Yang, Uiwon Hwang, Jaehyeong Sim
arXiv Computation and Language
Aug 24

When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA

arXiv:2511.17886v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved remarkable success across multimodal tasks, yet their substantial computational demands hinder ef...

By Pume Tuchinda, Parinthapat Pengpun, Romrawin Chumpu, Patomporn Payoungkhamdee, Sarana Nutanong, Peerat Limkonchotiwat
arXiv Computer Vision
Aug 27

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

MLLMCLIP introduces a heterogeneous distillation framework that transfers multimodal knowledge from a generative Multimodal Large Language Model (MLLM) teacher directly into a discriminative CLIP student, eliminating the need for synthetic hard negatives. The method uses an attention-based per-layer token selection and a CKA-based distillation loss to bridge architectural differences between the two models. As a result, MLLMCLIP achieves state‑of‑the‑art compositional accuracy and improves zero‑shot classification and image‑text retrieval performance.

By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
arXiv Machine Learning
Jul 28

Asymmetric Hierarchical Anchoring for Robust Audio-Visual Cross-Modal Generalization

arXiv:2602. 03570v2 Announce Type: replace Abstract: Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space.

By Bixing Wu, Yuhong Zhao, Zongli Ye, Jiachen Lian, Xiangyu Yue, Gopala Anumanchipalli
arXiv Computer Vision
2d ago

Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring

Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.

By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero