arXiv Computation and Language

Orthogonal Yet Coupled: Decoupling Geometric Components for Model Merging

The paper introduces DiGA, a Disentangled Geometry-Aware framework for merging pretrained models. DiGA orthogonally decomposes each task vector into components tied to distinct geometric attributes, aggregates these components independently, and then recombines them, thereby preserving each component’s geometric identity. Experiments across various models, tasks, and merging methods show that DiGA improves merged-model performance and reduces capability degradation.

arXiv Machine Learning
Aug 27

Escaping Low-Dimensional Overlap: Multi-Task Model Merging via High-Dimensional Sparse Disentanglement

The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.

By Yihang Zhang, Shengke Sun, Junjie Wen, Feng Zeng
arXiv Machine Learning
1d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
Sep 22

CAMFT: Conflict-Aware Mergeable Fine-Tuning for Large Language Models

CAMFT is a Conflict‑Aware Mergeable Fine‑Tuning method designed to make task adaptation efficient and merge‑aware for large language models. Unlike existing approaches that only resolve parameter conflicts after fine‑tuning, CAMFT shapes mergeability during training by guiding each task to update sparse coordinates with lower cross‑task conflict. Experiments show that CAMFT outperforms standard fine‑tuning baselines in multi‑task merging scenarios.

By Jingang Zhou, Haiyang Guo, Yuan Ma, Han Zhu, Xu-Yao Zhang
Hugging Face Trending Papers
Aug 18

CORAM: Coherent Orthogonal Rotation for Model Merging

CORAM introduces a new approach to merging fine‑tuned models by partitioning each target weight matrix into row slices and representing each slice with its singular value decomposition in the base‑model’s SVD frame. The method performs manifold averaging of task‑specific factors and applies an amplification coefficient to counteract contraction, with the coefficient’s scale estimated from update norms and its restoration strength chosen from expert update dispersion. Across multiple model families and scales, CORAM outperforms the prior OrthoMerge technique by up to 1.35 points and matches or exceeds the strongest weight‑space baselines.

arXiv Machine Learning
Aug 19

CORAM: Coherent Orthogonal Rotation for Model Merging

CORAM (Coherent Orthogonal Rotation for Model Merging) is a new method for combining fine‑tuned models without joint training or access to original data. It partitions each target weight matrix into row slices, represents each expert slice with its singular value decomposition in the base‑model SVD frame, and merges the task‑specific factors on their corresponding manifolds. The approach includes an amplification coefficient to counteract manifold averaging contraction, spread slicing to balance highly updated rows, and a residual pathway for non‑target layers, achieving improvements over existing orthogonal merging techniques across multiple model families and scales.

By Xinyi Sui, Ziran Liu, Nam Ling, Wei Wang, Wei Jiang