arXiv:2312. 06173v2 Announce Type: replace Abstract: Merging models fine-tuned from a common, extensively pre-trained large model but specialized for different tasks has been demonstrated as a cheap and scalable strategy to construct a multi-task model that performs well across diverse tasks.
By Anke Tang, Xianglin Luo, Li Shen, Yong Luo, Liang Ding, Han Hu, Bo Du, Dacheng Tao
arXiv:2606. 19164v1 Announce Type: cross Abstract: Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained checkpoint into a single model.
By Longhua Li, Lei Qi, Xin Geng, Qi Tian
The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.
By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
By Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao
arXiv:2606. 22589v2 Announce Type: replace Abstract: Ever since the advent of foundation models and the pre-training-finetuning paradigm, there have been numerous efforts to merge multiple task-specific experts into a single multi-task model.
By Jungyong Son, Jinwook Jung, Sungyong Baik
The paper proposes a sensitivity‑aware residual‑stream pruning method for large language models that goes beyond minimizing activation reconstruction error. By using a second‑order approximation of output KL divergence, the authors derive a spectral upper bound that selects pruning subspaces based on both activation covariance and output sensitivity, enabling efficient eigendecomposition. Experiments on instruction‑tuned language models show that this approach consistently reduces calibration KL divergence, improves perplexity, and enhances downstream task performance across various compression levels.
By Chayne Thrash, Kevin Chen, Soheil Kolouri