arXiv AI By Chao Wang, Yuchen Guo, Zheng Tan, Guanchun Wang, Yanbiao Ma, Qiqi Duan, Peng Wu

Model Merging to Evolution: Parameter Space Exploration for Expert Models

Read the original on arXiv AI →

arXiv:2606. 28373v1 Announce Type: cross Abstract: Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training, thereby reducing computational resource requirements.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian
arXiv Machine Learning
1d ago

Mixture-Trained Merging for Unified Multi-Objective Models

Mixture-Trained Merging (MTM) is a method for creating unified language models that combine multiple objectives—such as mathematics, code, instruction following, and controllable thinking—into a single parameter set. Instead of sequentially post‑training on each objective, MTM trains each branch on a mixture of objectives, ensuring that the branches remain compatible in weight space and can be merged without degrading performance. The approach iteratively refines merge coefficients using low‑cost evaluations and multi‑objective Bayesian optimization, outperforming naive merging and preserving distinct behaviors across domains.

By SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee
Hugging Face Trending Papers
Jun 25

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.