Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.
arXiv:2606. 22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model.
By Jungyong Son, Jinwook Jung, Sungyong Baik
arXiv:2606. 19164v1 Announce Type: cross Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint into a single model.
By Longhua Li, Lei Qi, Xin Geng, Qi Tian
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
By Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao
Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.
By Aditya Pola, Vineeth N. Balasubramanian
arXiv:2607. 01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.
By Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang
ReForge is a bilevel optimization framework that refines merged models by treating module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level produces a closed‑form MAP estimate from unlabeled calibration activations, while the outer level employs Bayesian optimization to jointly select regularization strengths and assembly scales using validation data. A data‑free variant replaces activation statistics with task‑vector Grams, enabling refinement without calibration examples, and across extensive vision and language benchmarks ReForge consistently outperforms existing plug‑and‑play anchor baselines, achieving significant accuracy gains on large‑scale tasks such as 20‑task ViT‑B/32 and eight‑task ViT‑L/14.
By Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji
arXiv:2606. 07500v1 Announce Type: cross Abstract: Continual learning in Large Language Models (LLMs) is hindered by the plasticity-stability dilemma, where acquiring new capabilities often leads to catastrophic forgetting of previous knowledge.
By Fatema Siddika, Md Anwar Hossen, Tanwi Mallick, Ali Jannesari
The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.
By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
arXiv:2507. 09471v4 Announce Type: replace Abstract: Continual Learning (CL) empowers AI models to continuously learn from sequential task streams.
By Lingfeng He, De Cheng, Zhiheng Ma, Huaijie Wang, Dingwen Zhang, Nannan Wang, Xinbo Gao
CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.
By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
arXiv:2506. 14126v2 Announce Type: replace-cross Abstract: Modern deep learning is increasingly characterized by the use of open-weight foundation models that can be fine-tuned on specialized datasets.
By Stefan Horoi, Guy Wolf, Eugene Belilovsky, Gintare Karolina Dziugaite