arXiv AI

A Generalization of Amari's Bayesian Duality

arXiv Machine Learning
Sep 18

Bayesian Optimization with Rich Auxiliary Information via LLMs

The paper introduces Bayesian Optimization (BO) techniques that incorporate rich auxiliary information—such as training curves, expert notes, images, and prior knowledge—using large language models (LLMs). Three new methods are proposed to integrate this auxiliary data into BO, and they are evaluated on hyperparameter optimization benchmarks and a real-world nuclear fusion task. The results show that these LLM-enhanced BO methods consistently outperform standard BO and existing LLM-based optimization approaches.

By Tejus Gupta, Efe Mert Karag\"ozl\"u, Rohit Sonker, Barnab\'as P\'oczos, Jeff Schnieder
arXiv Machine Learning
Jun 19

Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.

By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta
arXiv AI
Jul 7

Machine Unlearning via Information Theoretic Regularization

arXiv:2502. 05684v5 Announce Type: replace-cross Abstract: How can we effectively remove or ``unlearn'' undesirable information, such as specific features or the influence of individual data points, from a learning outcome while minimizing utility loss and ensuring rigorous guarantees?

By Shizhou Xu, Thomas Strohmer
arXiv Machine Learning
Jun 2

Symmetries in PAC-Bayesian Learning

arXiv:2510. 17303v2 Announce Type: replace Abstract: Symmetries are known to improve the empirical performance of machine learning models, yet theoretical guarantees explaining these gains remain limited.

By Armin Beck, Peter Ochs
arXiv Machine Learning
Jun 5

Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood Maximization

arXiv:2410. 02628v5 Announce Type: replace Abstract: Learning conditional distributions $\pi^*(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^*$.

By Mikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev, Dmitry Baranchuk, Anastasis Kratsios, Evgeny Burnaev, Alexander Korotin