arXiv Machine Learning By Yongxian Wei, Runxi Cheng, Xingxuan Zhang, Li Shen, Chun Yuan, Peng Cui, Dacheng Tao

Closed-Form Spectral Regularization for Multi-Task Model Merging

Read the original on arXiv Machine Learning →

arXiv:2606. 07289v1 Announce Type: new Abstract: Model merging combines several independently fine-tuned experts into a single multi-task model without any training data, reducing the storage, serving, and decentralized-development costs of large foundation models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion

arXiv:2312. 06173v2 Announce Type: replace Abstract: Merging models fine-tuned from a common, extensively pre-trained large model but specialized for different tasks has been demonstrated as a cheap and scalable strategy to construct a multi-task model that performs well across diverse tasks.

By Anke Tang, Xianglin Luo, Li Shen, Yong Luo, Liang Ding, Han Hu, Bo Du, Dacheng Tao
arXiv Machine Learning
1d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv Machine Learning
1d ago

Output-aware Residual Stream Pruning for Large Language Models

The paper proposes a sensitivity‑aware residual‑stream pruning method for large language models that goes beyond minimizing activation reconstruction error. By using a second‑order approximation of output KL divergence, the authors derive a spectral upper bound that selects pruning subspaces based on both activation covariance and output sensitivity, enabling efficient eigendecomposition. Experiments on instruction‑tuned language models show that this approach consistently reduces calibration KL divergence, improves perplexity, and enhances downstream task performance across various compression levels.

By Chayne Thrash, Kevin Chen, Soheil Kolouri