arXiv:2607. 01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.
By Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang
ReForge is a bilevel optimization framework that refines merged models by treating module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level produces a closed‑form MAP estimate from unlabeled calibration activations, while the outer level employs Bayesian optimization to jointly select regularization strengths and assembly scales using validation data. A data‑free variant replaces activation statistics with task‑vector Grams, enabling refinement without calibration examples, and across extensive vision and language benchmarks ReForge consistently outperforms existing plug‑and‑play anchor baselines, achieving significant accuracy gains on large‑scale tasks such as 20‑task ViT‑B/32 and eight‑task ViT‑L/14.
By Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji
The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.
By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv:2607. 17674v1 Announce Type: cross Abstract: A language model $p_\theta(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution.
By Awni Altabaa, John Lafferty
arXiv:2606. 01954v1 Announce Type: new Abstract: Implicit-process priors define distributions over functions through flexible generative mechanisms, making them attractive for Bayesian function-space modelling.
By Luis A. Ortega, Andr\'es R. Masegosa, Thomas D. Nielsen
arXiv:2607. 09073v1 Announce Type: new Abstract: Bayesian optimization routinely warm-starts a target experiment with data from related source tasks, and the multi-task Gaussian process is the textbook surrogate for the job.
By Carl Hvarfner, Sam Daulton, Max Balandat, Eytan Bakshy