arXiv Computer Vision By Pau de Jorge, C\'esar Roberto de Souza, Bj\"orn Michele, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Philippe Weinzaepfel, Florent Perronnin, Diane Larlus, Yannis Kalantidis

Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks

Read the original on arXiv Computer Vision →

The paper introduces the task alignment proxy, a method that accelerates hyperparameter selection for merging models fine‑tuned on diverse vision tasks. It addresses the challenge of training heterogeneous decoders, which makes traditional downstream performance evaluation costly. By using the proxy, the authors demonstrate that model merging can be applied efficiently to multi‑task vision models beyond CLIP‑based classification.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
1d ago

Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance

The paper introduces a method for selecting the best vision‑language model for a downstream task by analyzing the internal dynamics of the visual encoder. It represents each task with layer‑wise conductance and uses an entropy‑regularized alignment to derive a target‑conditioned block importance distribution. The proposed Directional Conductance Divergence (DCD) metric captures asymmetric transferability, enabling accurate prediction of model rankings without direct inference, and achieves a 14.7% NDCG@5 improvement over SWAB on 48 VLMs across 21 datasets.

By Wei Yang, Hong Xie, Tao Tan, Xin Li, Defu Lian, Enhong Chen
arXiv Machine Learning
Aug 10

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

arXiv:2608. 06901v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational and memory requirements pose significant challenges for deployment in constrained environments.

By Minseok Kang, Hyunwoo Kim, Chanyoung Kim, Minwoo Kim, Jaekoo Lee, Dahuin Jung