ReForge is a bilevel optimization framework that refines merged models by treating module-wise refinement as Bayesian linear regression with an anchor-centered prior. The inner level produces a closed‑form MAP estimate from unlabeled calibration activations, while the outer level employs Bayesian optimization to jointly select regularization strengths and assembly scales using validation data. A data‑free variant replaces activation statistics with task‑vector Grams, enabling refinement without calibration examples, and across extensive vision and language benchmarks ReForge consistently outperforms existing plug‑and‑play anchor baselines, achieving significant accuracy gains on large‑scale tasks such as 20‑task ViT‑B/32 and eight‑task ViT‑L/14.
By Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji
The paper introduces COVER, a multi‑task learning framework that regularizes covariate overlap to mitigate the negative effects of sharing information across tasks with differing covariate distributions and response relationships. COVER blends a common component function, a shared neural representation, and low‑dimensional task‑specific coefficients, using taskwise second‑moment matrices to guide coefficient integration. The authors provide theoretical bias‑variance analysis, oracle inequalities, and neural‑network convergence rates, and demonstrate that COVER outperforms existing deep‑learning and statistical integration methods in simulations and a GTEx central‑nervous‑system study.
By Yang Sui, Qi Xu, Yang Bai, Annie Qu
The paper introduces Gradient-based Sample Selection Bayesian Optimization (GSSBO), a method that builds the Gaussian process surrogate on a strategically chosen subset of samples rather than the full dataset. By using gradient information to eliminate redundant points while keeping diversity and representativeness, GSSBO achieves sublinear regret bounds and reduces the cubic computational cost of standard BO. Experiments on synthetic and real-world tasks show that this approach maintains comparable optimization performance while significantly cutting GP fitting time and resource usage.
By Qiyu Wei, Haowei Wang, Zirui Cao, Songhao Wang, Richard Allmendinger, Mauricio A \'Alvarez
arXiv:2607. 27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive.
By Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa
arXiv:2607. 01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.
By Long Minh Bui, Tuan Anh Le Van, Tung Phi Duc, Phi Le Nguyen, Jana Doppa, Trong Nghia Hoang
The paper presents Iterative Sequential Transfer (IST), a method for few-shot multiobjective multitask optimization that addresses the bottleneck of aligning elite solution distributions across tasks. IST treats multitask optimization as a sequence of transfer problems, focusing evaluations on one target task per iteration and using a likelihood-informed prioritization to select the task most ready for knowledge integration. Experiments on benchmark and real-world problems demonstrate IST’s effectiveness under tight evaluation budgets.