arXiv Computer Vision

Dyna-DINO: Efficient ViT Distillation Via Adaptive Representation Anchoring

Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.

arXiv Computer Vision
Aug 24

Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models

The paper introduces Semantic Localization-Enhanced Teacher (SLE‑T), a knowledge‑distillation framework that aligns spatial‑scale and semantic features between a Vision Foundation Model (VFM) teacher and a student detector for cross‑domain object detection. SLE‑T employs a lightweight SLE Adapter that injects pretrained local‑texture priors into DINOv2 and reformulates its features into dense, spatially and semantically compatible representations, enabling effective pseudo‑label learning or feature alignment. Experiments on three domain‑adaptive object detection benchmarks show that SLE‑T with DINOv2‑B achieves state‑of‑the‑art performance while using only a quarter of the training time and less GPU memory compared to the larger DINOv2‑G teacher.

By Qifeng Zhang, Ting Xiang, Zeyuan Bai, Changjian Chen
arXiv Computer Vision
Aug 28

Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification

The paper presents a method for cross‑architecture knowledge distillation from a fine‑tuned DINOv2 Vision Transformer teacher to a lightweight bidirectional Visual State Space Model (LVSSM) student for tea leaf disease classification. By addressing training‑stability issues with a progressive convolutional stem and gated selective‑scan block, the 4.45 M‑parameter student achieves a mean test accuracy of 95.41%—a 3.09‑point improvement over the teacher’s 92.32%—while using only one‑fifth of the teacher’s parameters. Ablation studies show that simple logit‑level distillation outperforms intermediate feature alignment, and the gains are specific to students that start below the teacher’s performance.

By Zibo Zhou, Zongsen Qiu, Rui Chen, Yujie Yao, Yue Zhou, Jianjun Wang
arXiv Machine Learning
2d ago

Decoupled and Distilled: Task-Adaptive LoRA-Teachers with Ensemble Knowledge Transfer for Few-Shot Class-Incremental Learning

The paper introduces TALON, a Task‑Adaptive LoRA‑Teacher framework for Few‑Shot Class‑Incremental Learning. TALON assigns a dedicated LoRA‑Teacher to each incremental task, then distills the frozen teachers into a single LoRA‑Student via Ensemble Knowledge Transfer, using a semantic‑guided weighting scheme to reduce forgetting and overfitting. Experiments on four FSCIL benchmarks show that TALON matches or surpasses state‑of‑the‑art accuracy while using up to 33× fewer deployment parameters and cutting inference time by 41.7%.

By Hongwei Zhao (School of Computer Science,Engineering, Beihang University), Rui Liu (School of Computer Science,Engineering, Beihang University), Yansong Liu (School of Computer Science,Engineering, Beihang University), Zhiyuan Zou (School of Computer Science,Engineering, Beihang University), Yong Chen (School of Computer Science, Beijing University of Posts,Telecommunications)
arXiv Machine Learning
Aug 4

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

arXiv:2604. 15622v3 Announce Type: replace-cross Abstract: Always-on contextual AI runs language-aligned vision foundation models (VFMs) on edge devices, where the on-device model is the dominant continuous compute cost under strict latency and power limits.

By Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li
arXiv Computer Vision
Aug 31

Token-Budget Distillation: Transferring Full-Token Semantics to Compressed Video Vision-Language Models

Token-Budget Distillation (TBD) is a parameter‑efficient fine‑tuning framework that adapts video vision‑language models to a fixed token budget. It freezes the pretrained backbone, updates only LoRA adapters, and incorporates FlashVID visual token compression. TBD uses a dual‑path teacher‑student design with full‑token supervision and compressed student optimization, enabling the student to recover full‑token semantics while remaining efficient under aggressive token reduction.

By Xiaoyang Guo, Guoping Luo, Jusheng Zhang, Keze Wang, Wenhao Wang
arXiv Computation and Language
Sep 10

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

arXiv:2609.10321v1 Announce Type: new Abstract: Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in curren...

By Hongyuan Zhang, Xianda Guo, Yanlun Peng, Qianlong Yang, Yubin Guo, Pinhan Fu, Mulin Chen, Xiaozhen Qiao, Ping Luo
arXiv Computation and Language
Aug 24

When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA

arXiv:2511.17886v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have achieved remarkable success across multimodal tasks, yet their substantial computational demands hinder ef...

By Pume Tuchinda, Parinthapat Pengpun, Romrawin Chumpu, Patomporn Payoungkhamdee, Sarana Nutanong, Peerat Limkonchotiwat