arXiv AI

Surrogate Benchmarks for Model Merging Optimization

arXiv:2509. 02555v2 Announce Type: replace-cross Abstract: Model merging techniques aim to integrate the abilities of multiple models into a single model.

arXiv Machine Learning
5d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv Machine Learning
5d ago

Mixture-Trained Merging for Unified Multi-Objective Models

Mixture-Trained Merging (MTM) is a method for creating unified language models that combine multiple objectives—such as mathematics, code, instruction following, and controllable thinking—into a single parameter set. Instead of sequentially post‑training on each objective, MTM trains each branch on a mixture of objectives, ensuring that the branches remain compatible in weight space and can be merged without degrading performance. The approach iteratively refines merge coefficients using low‑cost evaluations and multi‑objective Bayesian optimization, outperforming naive merging and preserving distinct behaviors across domains.

By SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee
arXiv AI
Sep 2

Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search

The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.

By Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
arXiv AI
Sep 3

CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging

CoMerge is a conflict‑driven preference optimization framework for merging multiple expert language models into a single multi‑task model without full retraining. It treats model merging as a preference optimization problem, using self‑supervised, conflict‑driven hard negative samples derived from naive merging defects to refine lightweight, tensor‑wise merging coefficients. Experiments show CoMerge achieves near‑perfect performance on MergeBench and improves instruction‑following and safety on Llama‑3.1‑8B‑Instruct while optimizing only 1,445 scalar coefficients.

By Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng