arXiv:2606. 10504v1 Announce Type: new Abstract: Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e.
By Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang
Cross-modal knowledge distillation (CMKD) studies how a (large) teacher model trained on one type of data (e. g.
CLIP-RD introduces a relational distillation framework for efficient CLIP knowledge distillation, featuring Vertical Relational Distillation (VRD) and Cross Relational Distillation (XRD). VRD aligns intra‑modal similarity distributions between teacher and student, while XRD aligns cross‑modal similarity distributions to enforce bidirectional symmetry. This joint modeling of multidirectional relational structures improves the student’s embedding geometry, yielding a 1.8%p performance gain over CLIP‑KD across various architectures, tasks, and corruption settings with minimal training‑time overhead.
By Jeannie Chung, Hanna Jang, Ingyeong Yang, Uiwon Hwang, Jaehyeong Sim
arXiv:2602. 17395v2 Announce Type: replace-cross Abstract: Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known classes.
By Lorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. Bagdanov
arXiv:2511.17886v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) have achieved remarkable success across multimodal tasks, yet their substantial computational demands hinder ef...
By Pume Tuchinda, Parinthapat Pengpun, Romrawin Chumpu, Patomporn Payoungkhamdee, Sarana Nutanong, Peerat Limkonchotiwat
MLLMCLIP introduces a heterogeneous distillation framework that transfers multimodal knowledge from a generative Multimodal Large Language Model (MLLM) teacher directly into a discriminative CLIP student, eliminating the need for synthetic hard negatives. The method uses an attention-based per-layer token selection and a CKA-based distillation loss to bridge architectural differences between the two models. As a result, MLLMCLIP achieves state‑of‑the‑art compositional accuracy and improves zero‑shot classification and image‑text retrieval performance.
By Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao, Hiromi Wakaki, Junmo Kim, Yuki Mitsufuji
Pretrained vision-language models such as CLIP excel at zero-shot recognition but often fail at compositionality, particularly attribute-object and relational structures. Recent studies mitigate this...
arXiv:2412. 01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks.
By Qianhan Feng, Wenshuo Li, Tong Lin, Xinghao Chen
arXiv:2602. 03570v2 Announce Type: replace Abstract: Audio-visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space.
By Bixing Wu, Yuhong Zhao, Zongli Ye, Jiachen Lian, Xiangyu Yue, Gopala Anumanchipalli
arXiv:2606. 03052v1 Announce Type: new Abstract: Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations remain underexplored.
By Seungu Kang, Songkuk Kim
arXiv:2609.16689v1 Announce Type: new
Abstract: Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained...
By Jinwoo Jeon, GyuYeop Do, Yubin Lim, Nam-Joon Kim, Hyun Gon Ryu, Hyuk-Jae Lee, Byung-Jun Lee
Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.
By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero