arXiv:2602. 12952v3 Announce Type: replace Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every model variant.
By Filippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello, Simone Calderara
arXiv:2606. 01503v1 Announce Type: cross Abstract: Unified vision-language models (VLMs) integrate visual understanding and visual generation within a single autoregressive backbone, but their joint training is computationally expensive and largely overlooked from an efficiency perspective.
By Siyi Chen, Weiming Zhuang, Jingtao Li, Lingjuan Lv
arXiv:2606. 22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model.
By Jungyong Son, Jinwook Jung, Sungyong Baik
The paper introduces a method for selecting the best vision‑language model for a downstream task by analyzing the internal dynamics of the visual encoder. It represents each task with layer‑wise conductance and uses an entropy‑regularized alignment to derive a target‑conditioned block importance distribution. The proposed Directional Conductance Divergence (DCD) metric captures asymmetric transferability, enabling accurate prediction of model rankings without direct inference, and achieves a 14.7% NDCG@5 improvement over SWAB on 48 VLMs across 21 datasets.
By Wei Yang, Hong Xie, Tao Tan, Xin Li, Defu Lian, Enhong Chen
arXiv:2510. 17426v3 Announce Type: replace-cross Abstract: The "alignment tax" of post-training is typically framed as a drop in task accuracy.
By Tiancheng Hu, Benjamin Minixhofer, Nigel Collier
arXiv:2608. 06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments.
By Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung