CORAM introduces a new approach to merging fine‑tuned models by partitioning each target weight matrix into row slices and representing each slice with its singular value decomposition in the base‑model’s SVD frame. The method performs manifold averaging of task‑specific factors and applies an amplification coefficient to counteract contraction, with the coefficient’s scale estimated from update norms and its restoration strength chosen from expert update dispersion. Across multiple model families and scales, CORAM outperforms the prior OrthoMerge technique by up to 1.35 points and matches or exceeds the strongest weight‑space baselines.
ResMerge is a new framework for merging large language models trained via reinforcement learning. It separates each model’s task vector into a leading spectral head and a residual component, finding that both parts contain valuable behavior knowledge but behave differently during merging. The method builds a stable residual backbone using Spherical Residual Consensus Adaptation and then adds a lightweight head correction module that activates only when experts agree, leading to better preservation of expert capabilities compared to existing merging baselines.
By Yandu Sun, Zhiyan Hou, Hongyan An, Weizhen Wang, Haokai Ma, Yuheng Jia, Junfeng Fang, Haiyun Guo, Jinqiao Wang
The paper introduces DiGA, a Disentangled Geometry-Aware framework for merging pretrained models. DiGA orthogonally decomposes each task vector into components tied to distinct geometric attributes, aggregates these components independently, and then recombines them, thereby preserving each component’s geometric identity. Experiments across various models, tasks, and merging methods show that DiGA improves merged-model performance and reduces capability degradation.
By Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie, Mengjie Zhao, Yunpu Ma, Kang Liu, Zihan Wang, Shi Feng, Daling Wang, Hinrich Sch\"utze
arXiv:2608. 07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices.
By Inesh Chakrabarti, Sourjya Roy, Bowen Bao, Thiago Crepaldi, Spandan Tiwari, Ashish Sirasao
arXiv:2606. 01717v1 Announce Type: new Abstract: Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered by gradient interference and bandwidth-heavy synchronization.
By Minsik Choi, Geewook Kim
arXiv:2606. 19549v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) makes it cheap to train many domain- and task-specific language model adapters, but whether two adapters can be merged is usually discovered only after both have been fully trained and evaluated.
By Lin Tang, Wei Zhang, Jing Li, Hongyu Chen, Ming Zhao, Yuxuan Wang
arXiv:2607. 20561v1 Announce Type: new Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter storage and task selection at inference time.
By Keumseo Ryum, Joonhyuk Kang
arXiv:2606. 03723v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) enables parameter-efficient specialization of foundation models, but the proliferation of task-specific adapters fragments capabilities across many adapters, complicating reuse and deployment.
By Zhengbao He, Ruiqi Ding, Zhehao Huang, Ruikai Yang, Tao Li, Xiaolin Huang
The paper introduces READ, a method for composing low‑rank adapters (LoRA) in large language models. By rewriting each adapter into a balanced canonical form and enforcing a one‑directional coupling, READ allows new skills to read but never write into the output subspaces of existing skills, eliminating interference. Experiments on four benchmark suites and two model families show that READ consistently outperforms existing baselines, improving SuperGLUE scores by over twenty points and domain suite scores by more than seven points.
By Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
arXiv:2607. 11997v1 Announce Type: cross Abstract: Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training.
By Nikita Kozodoi, Zainab Afolabi, Jack Butler
arXiv:2609.01129v1 Announce Type: new
Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under compos...
By Jiming Feng, Junliang Li
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
By Ningyuan Shi, Zhipeng Zhou, Hao Wang, Chunyan Miao, Peilin Zhao