arXiv Machine Learning

Saliency-Aware Model Merging

arXiv:2606. 00511v1 Announce Type: new Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domain proficiency.

arXiv Machine Learning
Sep 22

CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models

CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.

By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
arXiv Computer Vision
Sep 2

From Saliency to Discriminability: Rank-Preserving Visual Token Pruning for VLM Rerankers

arXiv:2609.00667v1 Announce Type: cross Abstract: Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning es...

By Siyi Liu, Hanjun Yang, Chenchen Zhang, Xiaorong Zhu, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
arXiv AI
Jul 14

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.

By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
arXiv Computer Vision
6d ago

Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation

The paper introduces Target Saliency Boosting, a new task that enhances the visual prominence of a specific object in text-to-image generation without visual priors. It proposes GazeME, a lightweight framework that inserts learnable marker tokens around object descriptions to indicate which objects to emphasize or suppress. By building a saliency-semantics dataset and using Saliency Prior Marker Activation, GazeME learns to adjust markers during training and automatically applies them at inference, effectively boosting target saliency while maintaining semantic alignment and image quality.

By Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao
arXiv Computer Vision
Aug 27

RAVE: Re-Allocating Visual Attention in Large Multimodal Models

RAVE (Re-Allocating Visual Attention) is a lightweight pair‑gating mechanism that adds a learned query‑key bias to pre‑softmax attention scores over visual keys, derived from pre‑RoPE query and key features. It requires no architectural changes to the backbone and can be trained end‑to‑end with the rest of the model. Across multiple multimodal benchmarks, RAVE improves standard attention by an average of 3 points, especially on perception‑intensive tasks such as multilingual OCR, chart understanding, document VQA, and scene text VQA.

By Xi Leng, Xinhong Ma, Ziqiang Dong, Feng Zhang, Xiaoying Tang, Yang Yang, Guanjun Jiang