arXiv:2606. 22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model.
By Jungyong Son, Jinwook Jung, Sungyong Baik
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
By Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao
The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.
By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
Model merging provides an efficient way to construct multi-task generalist models without additional training, but its performance often degrades under severe task interference. Task interference in m...
Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.
The paper introduces Net Utility, a data‑free metric for selecting which singular directions of low‑rank adapters (LoRAs) to keep when merging across tasks. By scoring each direction for task utility and interference, and then globally selecting the highest‑scoring directions under a total budget constraint, the method avoids the uniform‑budget assumption that hampers existing merging techniques. Experiments on vision and language tasks show that Net Utility‑based rank allocation yields about a 2% performance gain over other merging methods.
By Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi
CoMerge is a conflict‑driven preference optimization framework for merging multiple expert language models into a single multi‑task model without full retraining. It treats model merging as a preference optimization problem, using self‑supervised, conflict‑driven hard negative samples derived from naive merging defects to refine lightweight, tensor‑wise merging coefficients. Experiments show CoMerge achieves near‑perfect performance on MergeBench and improves instruction‑following and safety on Llama‑3.1‑8B‑Instruct while optimizing only 1,445 scalar coefficients.
By Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng
arXiv:2608. 12842v1 Announce Type: new Abstract: Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requiring additional retraining.
By Yuchen Liu, Zongzhen Yang, Binhang Qi, Hailong Sun, Xiang Gao
arXiv:2602. 07218v2 Announce Type: replace-cross Abstract: Adaptability has been regarded as a central feature in the foundation models, enabling them to effectively acclimate to unseen downstream tasks.
By Gagik Magakyan, Amirhossein Reisizadeh, Chanwoo Park, Pablo A. Parrilo, Asuman Ozdaglar
CoMerge introduces a conflict‑driven preference optimization framework for merging multi‑task large language models, reframing merging as a preference problem that uses self‑supervised hard negative samples derived from naive merging defects. By optimizing lightweight, tensor‑wise merging coefficients, the method mitigates parameter‑space conflicts while preserving task‑specific capabilities. Experiments show CoMerge achieves an average normalized performance of 0.9968 on MergeBench and improves conflict‑sensitive tasks on Llama‑3.1‑8B‑Instruct, outperforming both data‑free and data‑driven baselines while optimizing only 1,445 scalar coefficients.
arXiv:2606. 26902v1 Announce Type: new Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference.
By Jinwook Jung, Taegyu Kim, Kumju Jo, Sungyong Baik
arXiv:2609.24517v1 Announce Type: cross
Abstract: Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a single multi-task model without additional joi...
By Hyunjoong Cho, Jinhyeok Jang