arXiv AI By Utkarsh Agarwal, Vamshi Bonagiri, Raul Astudillo, Monojit Choudhury

Multi-Objective Bayesian Optimization for Model Merging

Read the original on arXiv AI →

arXiv:2608. 14264v1 Announce Type: cross Abstract: Model merging combines trained models directly in weight space, offering a compute-efficient alternative to additional fine-tuning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
1d ago

Mixture-Trained Merging for Unified Multi-Objective Models

Mixture-Trained Merging (MTM) is a method for creating unified language models that combine multiple objectives—such as mathematics, code, instruction following, and controllable thinking—into a single parameter set. Instead of sequentially post‑training on each objective, MTM trains each branch on a mixture of objectives, ensuring that the branches remain compatible in weight space and can be merged without degrading performance. The approach iteratively refines merge coefficients using low‑cost evaluations and multi‑objective Bayesian optimization, outperforming naive merging and preserving distinct behaviors across domains.

By SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee
arXiv AI
Jul 13

LLM-Driven Evolutionary Generation of Multi-Objective Bayesian Optimization Algorithms

arXiv:2607. 08791v1 Announce Type: cross Abstract: Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdependent design choices whose optimal configuration is problem-dependent and typically demands deep expertise.

By Georgios Laskaris, Reuben Brasher, Niki van Stein, Elena Raponi, Thomas B\"ack, Florian Neukart
arXiv Machine Learning
Sep 24

tidyHEBO: Robust General-Purpose Bayesian Optimization with Model-Consistent Warping and Pareto Search

tidyHEBO is a BoTorch-native Bayesian optimization tool that jointly applies Yeo-Johnson output warping to a Gaussian‑process surrogate, evaluates acquisition functions on the original objective scale, and conducts constrained cumulative Pareto search across multiple acquisition criteria. Using only default settings, it outperformed other methods on the Olympus benchmark and performed strongly on synthetic, Needle‑in‑a‑Haystack, and Bayesmark tasks, while adaptive batching offered a trade‑off between parallelization and optimization quality. These results position tidyHEBO as a robust, reproducible optimizer suitable for diverse practical problems, including scientific applications and hyperparameter tuning.

By L. A. Zhukov, E. V. Shaburova, D. V. Antonets