arXiv Machine Learning

Closed-Form Spectral Regularization for Multi-Task Model Merging

arXiv:2606. 07289v1 Announce Type: new Abstract: Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-development costs of large foundation models.

arXiv Machine Learning
Jun 16

Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion

arXiv:2312. 06173v2 Announce Type: replace Abstract: Merging models fine-tuned from a common, extensively pre-trained large model but specialized for different tasks has been demonstrated as a cheap and scalable strategy to construct a multi-task model that performs well across diverse tasks.

By Anke Tang, Xianglin Luo, Li Shen, Yong Luo, Liang Ding, Han Hu, Bo Du, Dacheng Tao
arXiv Machine Learning
1d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv Machine Learning
1d ago

Output-aware Residual Stream Pruning for Large Language Models

The paper proposes a sensitivity‑aware residual‑stream pruning method for large language models that goes beyond minimizing activation reconstruction error. By using a second‑order approximation of output KL divergence, the authors derive a spectral upper bound that selects pruning subspaces based on both activation covariance and output sensitivity, enabling efficient eigendecomposition. Experiments on instruction‑tuned language models show that this approach consistently reduces calibration KL divergence, improves perplexity, and enhances downstream task performance across various compression levels.

By Chayne Thrash, Kevin Chen, Soheil Kolouri
arXiv Machine Learning
Aug 27

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.

By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
arXiv AI
Jul 7

Learning to Discover Iterative Spectral Algorithms

arXiv:2602. 09530v2 Announce Type: replace-cross Abstract: We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization.

By Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael W. Mahoney
arXiv Machine Learning
Jun 2

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models

arXiv:2606. 00573v1 Announce Type: new Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices.

By Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren, Sai Qian Zhang