arXiv AI

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

arXiv:2412. 08147v2 Announce Type: replace-cross Abstract: Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute.

arXiv AI
2d ago

ReForge: Refining Merged Models with Anchor-Regularized Regression

ReForge is a bilevel optimization framework that refines merged models by treating module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level produces a closed‑form MAP estimate from unlabeled calibration activations, while the outer level employs Bayesian optimization to jointly select regularization strengths and assembly scales using validation data. A data‑free variant replaces activation statistics with task‑vector Grams, enabling refinement without calibration examples, and across extensive vision and language benchmarks ReForge consistently outperforms existing plug‑and‑play anchor baselines, achieving significant accuracy gains on large‑scale tasks such as 20‑task ViT‑B/32 and eight‑task ViT‑L/14.

By Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji
arXiv Machine Learning
1d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau
arXiv Machine Learning
Aug 24

Exact and general decoupled solutions of the LMC Multitask Gaussian Process model

The paper presents an exact, efficient solution for the Linear Model of Co‑regionalization (LMC) multitask Gaussian Process by decoupling latent processes under a mild noise‑model assumption. It introduces a full parametrization of the resulting projected LMC, enabling linear‑time optimization and simplifying tasks such as training updates and leave‑one‑out cross‑validation. Experiments on synthetic and real data demonstrate that projected LMC is competitive with state‑of‑the‑art multitask GP models while offering greater interpretability and computational ease.

By Olivier Truffinet (CEA Saclay), Karim Ammar (CEA Saclay), Jean-Philippe Argaud (EDF R&D), Bertrand Bouriquet (EDF)
arXiv Machine Learning
Jul 30

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

arXiv:2607. 26924v1 Announce Type: new Abstract: Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse.

By Chang Liu, Fei Suo, Yanzhou Jin, Yusuke Iwasawa, Yutaka Matsuo, Yaonan Zhu
Hugging Face Trending Papers
Jul 29

Temporally Centered SIGReg Improves Multi-Task LeWorldModel Learning: From Analysis to Method

Recent work on LeWorldModel (LeWM) has shown that the Sketched Isotropic Gaussian Regularizer (SIGReg) enables stable end-to-end world-model learning from pixels by regularizing the latent marginal distribution toward an isotropic Gaussian, thereby preventing representation collapse. While effective and elegant in single-task settings, this recipe does not extend reliably to multi-task training, leading to substantially worse downstream behavior-cloning performance.