arXiv AI

c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperparameter Optimization

arXiv:2211. 14411v5 Announce Type: replace-cross Abstract: Hyperparameter optimization (HPO) is crucial for strong performance of deep learning algorithms and real-world applications often impose some constraints, such as on memory usage or latency, on top of the performance requirement.

arXiv Machine Learning
Sep 18

Bayesian Optimization with Rich Auxiliary Information via LLMs

The paper introduces Bayesian Optimization (BO) techniques that incorporate rich auxiliary information—such as training curves, expert notes, images, and prior knowledge—using large language models (LLMs). Three new methods are proposed to integrate this auxiliary data into BO, and they are evaluated on hyperparameter optimization benchmarks and a real-world nuclear fusion task. The results show that these LLM-enhanced BO methods consistently outperform standard BO and existing LLM-based optimization approaches.

By Tejus Gupta, Efe Mert Karag\"ozl\"u, Rohit Sonker, Barnab\'as P\'oczos, Jeff Schnieder
arXiv Machine Learning
Sep 15

Exploring new directions in enhancing the ACTS parameter optimization suite

The article investigates how Bayesian optimization can improve the ACTS parameter optimization suite for charged‑particle reconstruction. By comparing Expected Improvement and Upper Confidence Bound with TPE and random search on an eight‑parameter problem, extending the best method to fifteen parameters, and applying Expected Hypervolume Improvement for multi‑objective tuning, the study shows that Bayesian acquisition methods find strong configurations earlier and maintain advantages in held‑out validation. The results demonstrate that Bayesian optimization enhances ACTS auto‑tuning through more efficient evaluations, broader search spaces, and the ability to select from non‑dominated trade‑off solutions.

By Chance LaVoie, Qi Bin Lei, Rocky Bala Garg, Lauren Tompkins
arXiv AI
Aug 20

Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization

The paper presents a method that uses Constrained Bayesian Optimization (CBO) to minimize the energy consumption of machine learning models while ensuring their generalization performance stays above a specified threshold. By treating energy usage as the primary objective and performance as a constraint, the authors demonstrate that CBO can reduce training energy costs on both regression and classification tasks without sacrificing predictive accuracy.

By Pallavi Mitra, Felix Biessmann
arXiv Machine Learning
Sep 24

tidyHEBO: Robust General-Purpose Bayesian Optimization with Model-Consistent Warping and Pareto Search

tidyHEBO is a BoTorch-native Bayesian optimization tool that jointly applies Yeo-Johnson output warping to a Gaussian‑process surrogate, evaluates acquisition functions on the original objective scale, and conducts constrained cumulative Pareto search across multiple acquisition criteria. Using only default settings, it outperformed other methods on the Olympus benchmark and performed strongly on synthetic, Needle‑in‑a‑Haystack, and Bayesmark tasks, while adaptive batching offered a trade‑off between parallelization and optimization quality. These results position tidyHEBO as a robust, reproducible optimizer suitable for diverse practical problems, including scientific applications and hyperparameter tuning.

By L. A. Zhukov, E. V. Shaburova, D. V. Antonets
arXiv Machine Learning
Jul 31

Towards Stability of Parameter-Free Optimization

arXiv:2405. 04376v4 Announce Type: replace Abstract: Hyperparameter tuning, particularly the selection of an appropriate learning rate in adaptive gradient training methods, remains a challenge.

By Yijiang Pang, Shuyang Yu, Bao Hoang, Jiayu Zhou
arXiv Machine Learning
Aug 27

GRAPE: Gradient Refinement and Progress-Aware Exploitation for Query-Efficient High-Dimensional Bayesian Optimization

GRAPE is a two‑stage Bayesian optimization framework that first refines the local gradient posterior using a closed‑form acquisition function and then selects update directions by maximizing expected decrease conditioned on descent. The authors prove that the refinement stage monotonically reduces local uncertainty and that the progress‑aware direction converges to true steepest descent as the posterior sharpens. Empirical results show GRAPE achieves a 5.4× speedup on black‑box adversarial attacks and reduces final average regret by 3.8 log‑units on large language model prompt‑optimization tasks.

By Richard Cornelius Suwandi, Feng Yin