arXiv Machine Learning By Kuangpu Guo, Aijing Yu, Jian Liang, Yuhe Ding, Zilei Wang, Ran He, Tieniu Tan

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

Read the original on arXiv Machine Learning →

arXiv:2512. 01461v2 Announce Type: replace Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.

By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
Hugging Face Trending Papers
Jun 25

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.

arXiv Machine Learning
Sep 22

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

The paper introduces Net Utility, a data‑free metric for selecting which singular directions of low‑rank adapters (LoRAs) to keep when merging across tasks. By scoring each direction for task utility and interference, and then globally selecting the highest‑scoring directions under a total budget constraint, the method avoids the uniform‑budget assumption that hampers existing merging techniques. Experiments on vision and language tasks show that Net Utility‑based rank allocation yields about a 2% performance gain over other merging methods.

By Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi