arXiv Machine Learning

Accelerating Evolutionary Strategy via Rao-Blackwellizing Realization of Uncertain Input

arXiv:2608. 02073v1 Announce Type: cross Abstract: We investigate Optimization under Input Uncertainty (OIU), in which the input to the objective function, rather than the objective function itself, is subject to uncertainty.

arXiv Machine Learning
Sep 11

Fisher-Rao Gradient Flows of Linear Programs and State-Action Natural Policy Gradients

The paper investigates a natural gradient method based on the Fisher information matrix of state-action distributions, which follows a Fisher‑Rao gradient flow within the state-action polytope under a linear potential. It establishes linear convergence rates for Fisher‑Rao gradient flows of linear programs, with the rate tied to the program’s geometry, and provides improved error bounds for entropic regularization. Additionally, the authors extend their analysis to perturbed flows, proving sublinear convergence for both perturbed Fisher‑Rao and natural gradient flows, thereby encompassing state‑action natural policy gradients.

By Johannes M\"uller, Semih \c{C}ayc{\i}, Guido Mont\'ufar
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas
arXiv Machine Learning
Sep 18

Expected Hypervolume Maximization for Multiobjective Optimization under Uncertainties

The paper proposes a Bayesian decision framework for multiobjective optimization under uncertainty, focusing on maximizing the expected hypervolume over a finite set of input points. It demonstrates that gradient‑based stochastic optimization can be applied, especially when dominated points are handled carefully, and suggests using Gaussian Processes as differentiable surrogate models when direct gradients are unavailable. Additionally, the authors introduce active learning strategies via acquisition functions to build surrogate models tailored to the multiobjective problem and evaluate these strategies on simple analytical benchmarks.

By Victor Trappler (Mines Saint-\'Etienne MSE, LIMOS, FAYOL-ENSMSE, FAYOL-ENSMSE)