arXiv AI By Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

Read the original on arXiv AI →

arXiv:2607. 01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 25

Learning to Recover Task Experts from a Multi-Task Merged Model

Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers from parameter interference. While dynamic merging models aim to bridge this gap, many works rely on the costly storage and loading of redundant expert components at inference.

arXiv AI
2d ago

ReForge: Refining Merged Models with Anchor-Regularized Regression

ReForge is a bilevel optimization framework that refines merged models by treating module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level produces a closed‑form MAP estimate from unlabeled calibration activations, while the outer level employs Bayesian optimization to jointly select regularization strengths and assembly scales using validation data. A data‑free variant replaces activation statistics with task‑vector Grams, enabling refinement without calibration examples, and across extensive vision and language benchmarks ReForge consistently outperforms existing plug‑and‑play anchor baselines, achieving significant accuracy gains on large‑scale tasks such as 20‑task ViT‑B/32 and eight‑task ViT‑L/14.

By Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji
arXiv Machine Learning
Sep 22

Merge++: Universal Merge Refinement Through Data-Free Checkpoint Inversion

Merge++ is a post‑hoc refinement technique for model merging that synthesizes task‑representative images by inverting expert checkpoints and then distills expert knowledge into a merged model. It operates without any additional data beyond the checkpoints and can be applied universally across existing weight‑space merging algorithms. Experiments show consistent improvements, with average gains of +2 to +8 points and up to +25.9 on specific configurations.

By Aditya Pola, Vineeth N. Balasubramanian