arXiv:2608. 13596v1 Announce Type: cross Abstract: Heterogeneous model fusion seeks to combine models that differ in tasks, initializations, architectures, or scales.
By Jiahe Fan, Si Chen, Yinghao Hou, Aiyuan Zhang, Hong Xie
arXiv:2607. 18026v1 Announce Type: new Abstract: Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment?
By Jiahe Fan, Yinghao Hou, Si Chen, Aiyuan Zhang, Hong Xie, Defu Lian
arXiv:2602. 12952v3 Announce Type: replace Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every model variant.
By Filippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello, Simone Calderara
CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.
By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
The paper introduces Align‑LoRA, a unified LoRA framework for multi‑task learning that replaces complex, isolated adapter designs with a single‑adapter model enhanced by a higher rank and an explicit alignment loss. It demonstrates that a router‑free, multi‑head model with high inter‑head redundancy can outperform more elaborate baselines, and that a unified LoRA can achieve competitive performance while enabling weight merging and zero inference latency. Extensive experiments and theoretical analysis confirm that Align‑LoRA surpasses prevailing approaches, offering a simpler, production‑friendly paradigm for parameter‑efficient fine‑tuning of large language models.
By Jinda Liu, Yi Chang, Yuan Wu
The paper investigates three fusion paradigms—Merge, Mix RL, and multi‑teacher on‑policy distillation (MOPD)—for consolidating reinforcement learning with verifiable rewards (RLVR) across multiple domains. Experiments across model scales and a multi‑domain benchmark show that while overall performance differences are small, significant gaps can appear on specific tasks, and each method exhibits distinct training dynamics and constraints. Practical guidelines are offered: Merge for cheap fusion when experts exist, Mix RL for unified training with adjustable domain mixtures, and MOPD when preserving domain‑specific gains is paramount.
By Siye Wu, Kai Yang, Yuchen Cai, Xin Xu, Peng-Yuan Wang, Jiaxuan Wang, Jiashun Liu, Jiafei Lyu, Yangkun Chen, Saiyong Yang, Yanghua Xiao
The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.
By Baban Gain, Trilok Nath Singh, Asif Ekbal
The study investigates how fine‑tuning a large language model on a single task‑language pair influences performance on other task‑language pairs. Using LoRA fine‑tuning across multiple open‑weight LLM families, 11 languages, and four benchmarks, the authors decompose transfer into matched‑task, cross‑task, and cross‑task cross‑language regimes. They find that while single‑source fine‑tuning generally improves performance, the gains are highly asymmetric, with matched‑task cross‑language transfer being most effective and driven mainly by the target language rather than model architecture.
By Kajetan Dymkiewicz, Ivan Vulic, Helen Yannakoudakis, Eilam Shapira, Roi Reichart, Anna Korhonen
The paper proposes a method called Modular Expert Merging for Biomedical Retrieval, which combines independently trained domain‑specialized experts instead of large mixed‑domain training. Experiments across four decoder‑only LLM families (0.6B‑7B) and twelve retrieval tasks from MTEB show that merging experts consistently outperforms mixed‑domain training. The authors also introduce a Synthesize‑Train‑Merge (STM) framework that generates hard negatives with a top‑tier LLM, fine‑tunes experts via LoRA, and merges them, achieving strong biomedical retrieval performance while retaining competitive general‑domain results.
By Sameh Khattab, Jean-Philippe Corbeil, Osman Alperen \c{C}inar-Kora\c{s}, Amin Dada, Julian Friedrich, Jiawei He, Douglas Teodoro, Jens Kleesiek
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
By Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao
arXiv:2608. 09201v1 Announce Type: new Abstract: Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in weight space.
By Lingching Tung, Chi-Jui Kim, Beicheng Xu, Yuchen Wang, Bin Cui
arXiv:2608. 05980v1 Announce Type: new Abstract: We investigate whether simple transformations can translate representations across heterogeneous text embedding models.
By Sid Ali Hamideche (Orange Research), Louis Adrien Dufrene (Orange Research), Quentin Lampin (Orange Research), Guillaume Larue (Orange Research)