CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.
By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
arXiv:2604. 00757v2 Announce Type: replace-cross Abstract: Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens.
By Dong-Jae Lee, Sunghyun Baek, Junmo Kim
arXiv:2609.13706v1 Announce Type: new
Abstract: Co-salient object detection (Co-SOD) requires a model to find foreground regions that are salient in individual images and supported by the image group...
By Yuan Xiang, Matteo Rossi, Yingzhou Chen
arXiv:2609.37581v1 Announce Type: cross
Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visu...
By Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li
arXiv:2609.36916v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention...
By Weixuan Li, Zikun Zhou, Xinyi Zhuang, Xinyan Guo, Rui Tian, Chuyao Zhang, Lin Gao
arXiv:2609.00667v1 Announce Type: cross
Abstract: Large vision-language models used as listwise rerankers must jointly process visual tokens from tens of candidates per query, making token pruning es...
By Siyi Liu, Hanjun Yang, Chenchen Zhang, Xiaorong Zhu, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
arXiv:2608.06411v2 Announce Type: replace-cross
Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by...
By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang, Hao Geng, Minjun Yu
arXiv:2511. 11421v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) aims to continually learn new categories without forgetting previously acquired knowledge.
By Lan Li, Tao Hu, Da-Wei Zhou, Jia-Qi Yang, Han-Jia Ye, De-Chuan Zhan
arXiv:2608. 06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens.
By Yuyao Sun, Tao Deng, Shuang Li, Deqing Wang, Hao Geng, Minjun Yu
arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.
By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu
The paper introduces Target Saliency Boosting, a new task that enhances the visual prominence of a specific object in text-to-image generation without visual priors. It proposes GazeME, a lightweight framework that inserts learnable marker tokens around object descriptions to indicate which objects to emphasize or suppress. By building a saliency-semantics dataset and using Saliency Prior Marker Activation, GazeME learns to adjust markers during training and automatically applies them at inference, effectively boosting target saliency while maintaining semantic alignment and image quality.
By Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao
RAVE (Re-Allocating Visual Attention) is a lightweight pair‑gating mechanism that adds a learned query‑key bias to pre‑softmax attention scores over visual keys, derived from pre‑RoPE query and key features. It requires no architectural changes to the backbone and can be trained end‑to‑end with the rest of the model. Across multiple multimodal benchmarks, RAVE improves standard attention by an average of 3 points, especially on perception‑intensive tasks such as multilingual OCR, chart understanding, document VQA, and scene text VQA.
By Xi Leng, Xinhong Ma, Ziqiang Dong, Feng Zhang, Xiaoying Tang, Yang Yang, Guanjun Jiang