arXiv:2607. 01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.
By Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang
Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.
By Aditya Pola, Vineeth N. Balasubramanian
arXiv:2509. 02555v2 Announce Type: replace-cross Abstract: Model merging techniques aim to integrate the abilities of multiple models into a single model.
By Rio Akizuki, Yuya Kudo, Nozomu Yoshinari, Yoichi Hirose, Toshiyuki Nishimoto, Kento Uchida, Shinichi Shirakawa
arXiv:2606. 26902v1 Announce Type: new Abstract: Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference.
By Jinwook Jung, Taegyu Kim, Kumju Jo, Sungyong Baik
Mixture-Trained Merging (MTM) is a method for creating unified language models that combine multiple objectives—such as mathematics, code, instruction following, and controllable thinking—into a single parameter set. Instead of sequentially post‑training on each objective, MTM trains each branch on a mixture of objectives, ensuring that the branches remain compatible in weight space and can be merged without degrading performance. The approach iteratively refines merge coefficients using low‑cost evaluations and multi‑objective Bayesian optimization, outperforming naive merging and preserving distinct behaviors across domains.
By SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee
Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.