arXiv:2606. 01557v1 Announce Type: new Abstract: Everywhere learning is a new paradigm whereby Artificial Intelligence (AI) systems are trained to satisfy loss constraints with probability one over the data distribution.
By Ignacio Boero, Ignacio Hounie, Luiz Chamon, Alejandro Ribeiro
arXiv:2603. 09793v2 Announce Type: replace Abstract: Bayesian optimization is a data-efficient technique that has been shown to be extremely powerful to optimize expensive, black-box, and possibly noisy objective functions.
By Federico Pavesi, Antonio Candelieri, No\'emie Jaquier
arXiv:2609.36310v1 Announce Type: new
Abstract: Everywhere learning provides a principled framework for training AI models under constraints that must hold throughout the data distribution. In the du...
By Ignacio Boero, Jonathan Nixon, Alejandro Ribeiro
arXiv:2509. 21725v3 Announce Type: replace Abstract: A bilevel optimization problem consists of two optimization problems nested as an upper- and a lower-level problem, in which the optimality of the lower-level problem defines a constraint for the upper-level problem.
By Takuya Kanayama, Yuki Ito, Tomoyuki Tamura, Masayuki Karasuyama
The paper introduces Bayesian Optimization (BO) techniques that incorporate rich auxiliary information—such as training curves, expert notes, images, and prior knowledge—using large language models (LLMs). Three new methods are proposed to integrate this auxiliary data into BO, and they are evaluated on hyperparameter optimization benchmarks and a real-world nuclear fusion task. The results show that these LLM-enhanced BO methods consistently outperform standard BO and existing LLM-based optimization approaches.
By Tejus Gupta, Efe Mert Karag\"ozl\"u, Rohit Sonker, Barnab\'as P\'oczos, Jeff Schnieder
arXiv:2606. 13984v1 Announce Type: cross Abstract: Decision trees are one of the fundamental tools in statistical learning due to their interpretability, flexibility, and their ability to adapt to nonlinear structures.
By Mathias Bourel
arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.
By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta
arXiv:2502. 05684v5 Announce Type: replace-cross Abstract: How can we effectively remove or ``unlearn'' undesirable information, such as specific features or the influence of individual data points, from a learning outcome while minimizing utility loss and ensuring rigorous guarantees?
By Shizhou Xu, Thomas Strohmer
arXiv:2602. 23672v2 Announce Type: replace-cross Abstract: This study proposes a General Bayes framework for policy learning.
By Masahiro Kato
arXiv:2606. 05380v1 Announce Type: cross Abstract: We present learning-augmented algorithms for two general classes of online minimization problems: metrical task systems and laminar set cover.
By Christian Coester, Alexa Tudose, Alexander Turoczy
arXiv:2510. 17303v2 Announce Type: replace Abstract: Symmetries are known to improve the empirical performance of machine learning models, yet theoretical guarantees explaining these gains remain limited.
By Armin Beck, Peter Ochs
arXiv:2410. 02628v5 Announce Type: replace Abstract: Learning conditional distributions $\pi^*(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^*$.
By Mikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev, Dmitry Baranchuk, Anastasis Kratsios, Evgeny Burnaev, Alexander Korotin