arXiv Machine Learning

CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging

arXiv:2607. 20561v1 Announce Type: new Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter storage and task selection at inference time.

arXiv Machine Learning
Sep 22

Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks

The paper introduces Net Utility, a data‑free metric for selecting which singular directions of low‑rank adapters (LoRAs) to keep when merging across tasks. By scoring each direction for task utility and interference, and then globally selecting the highest‑scoring directions under a total budget constraint, the method avoids the uniform‑budget assumption that hampers existing merging techniques. Experiments on vision and language tasks show that Net Utility‑based rank allocation yields about a 2% performance gain over other merging methods.

By Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi
arXiv Machine Learning
Sep 22

CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models

CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.

By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
Jun 2

Saliency-Aware Model Merging

arXiv:2606. 00511v1 Announce Type: new Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domain proficiency.

By Jungin Park, Jiyoung Lee, Kwanghoon Sohn
Hugging Face Trending Papers
Jun 25

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.