arXiv Machine Learning

Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging

arXiv Computer Vision
Sep 22

Beyond Uniform Subspaces: Spectrum-Aware and Depth-Adaptive Fusion for Multi-Task Model Merging

The paper introduces SADA-Merging, a new data‑free model merging framework that addresses limitations of existing subspace‑based methods. It allocates subspace capacity based on each task’s spectral complexity, adapts spectral preservation to task‑specific plasticity, and applies depth‑dependent anchoring to mitigate projection distortion. Experiments show that SADA‑Merging consistently outperforms current data‑free merging techniques across various task scales and adaptation settings.

By Ruxi Gu, Zilei Wang, Wei Wang
arXiv Machine Learning
Sep 22

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

The paper introduces Net Utility, a data‑free metric for selecting which singular directions of low‑rank adapters (LoRAs) to keep when merging across tasks. By scoring each direction for task utility and interference, and then globally selecting the highest‑scoring directions under a total budget constraint, the method avoids the uniform‑budget assumption that hampers existing merging techniques. Experiments on vision and language tasks show that Net Utility‑based rank allocation yields about a 2% performance gain over other merging methods.

By Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi
arXiv Computation and Language
Aug 27

ResMerge: Residual-based Spectral Merging of Large Language Models

ResMerge is a new framework for merging large language models trained via reinforcement learning. It separates each model’s task vector into a leading spectral head and a residual component, finding that both parts contain valuable behavior knowledge but behave differently during merging. The method builds a stable residual backbone using Spherical Residual Consensus Adaptation and then adds a lightweight head correction module that activates only when experts agree, leading to better preservation of expert capabilities compared to existing merging baselines.

By Yandu Sun, Zhiyan Hou, Hongyan An, Weizhen Wang, Haokai Ma, Yuheng Jia, Junfeng Fang, Haiyun Guo, Jinqiao Wang