arXiv Machine Learning By Miroslav Krstic, Luke Bhan

Families of Control-Cost-Parametrized Inverse-Optimal Universal Stabilizers

Read the original on arXiv Machine Learning →

arXiv:2606. 09047v1 Announce Type: cross Abstract: A classical universal stabilization formula offers the practitioner no design freedom: it is a single, parameter-free object.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
4d ago

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He