arXiv Machine Learning

Information Theoretic Bayesian Optimization over the Probability Simplex

arXiv:2603. 09793v2 Announce Type: replace Abstract: Bayesian optimization is a data-efficient technique that has been shown to be extremely powerful to optimize expensive, black-box, and possibly noisy objective functions.

arXiv AI
6d ago

Bayesian Optimization with Fisher Information Geometry: Gradient Bounds and Trust-Region Methods

The paper investigates Bayesian optimization using information geometry, deriving a local sensitivity tensor from the Fisher information metric that bounds the gradient of reparameterizable acquisition functions. This framework explains vanishing-gradient issues in high-dimensional settings and unifies heuristics like RAASP and dimension-scaled lengthscales. Leveraging this insight, the authors introduce FITR, a trust‑region BO method that replaces lengthscale scaling with local pullback‑Fisher weights, achieving competitive performance on GP benchmarks and extending naturally to non‑isotropic surrogates.

By Saksham Kiroriwal, Julius Pfrommer, J\"urgen Beyerer
arXiv Machine Learning
Jun 19

Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.

By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta
arXiv Machine Learning
Jun 9

Improving Bayesian Optimization via Training-Aware Conditional Diffusion Models

arXiv:2606. 08438v1 Announce Type: cross Abstract: Bayesian optimization (BO) is a widely used approach for black-box optimization that uses a Gaussian process (GP) as a surrogate and guides sequential evaluations via an acquisition function, with the ultimate goal of locating the global optimum $\mathbf{x}^{\star}$.

By Yilin Zheng, Haowei Wang, Szu Hui Ng, Enlu Zhou
arXiv Machine Learning
Aug 18

Iso-Riemannian Optimization on Learned Data Manifolds

arXiv:2510. 21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited.

By Willem Diepeveen, Melanie Weber
arXiv AI
Sep 10

A Generalization of Amari's Bayesian Duality

arXiv:2609.09126v1 Announce Type: new Abstract: Amari's contributions to information geometry and machine learning are well known. Here, we revisit Amari's work on Bayesian duality which has not rece...

By Mohammad Emtiyaz Khan, Thomas M\"ollenhoff
arXiv Machine Learning
Sep 4

No-Regret Bayesian Optimization with Finite-Library Input-Warped Kernels

The paper introduces Finite-Library Input-Warped Bayesian Optimization (FLIWBO), a method that selects input warps from a finite library to adapt the geometry used by Gaussian‑process Bayesian optimization. FLIWBO maintains high‑probability convergence guarantees while improving sample efficiency on problems where raw coordinates poorly match the objective’s geometry, such as log‑scaled hyperparameters or localized peaks. Experiments on synthetic benchmarks, Fashion‑MNIST hyperparameter tuning, and a 20‑dimensional multi‑agent system design demonstrate that FLIWBO‑UCB outperforms raw‑coordinate GP‑UCB and other methods with regret guarantees, especially under misspecified geometry.

By Edvin Ketabati Augustinsson, Robert A. Bridges