arXiv Machine Learning

Families of Control-Cost-Parametrized Inverse-Optimal Universal Stabilizers

arXiv:2606. 09047v1 Announce Type: cross Abstract: A classical universal stabilization formula offers the practitioner no design freedom: it is a single, parameter-free object.

arXiv Machine Learning
4d ago

Global Optimality for Constrained Exploration via Penalty Regularization

The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.

By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv Machine Learning
Sep 17

Stability-Constrained Approximation in Spline KANs: Exact Layer Balancing and Budget-Compatible Saturation

The paper investigates how to balance approximation accuracy and stability in deep spline superposition networks under a strict layerwise Lipschitz budget. It provides an exact solution to the finite‑depth diagonal balancing problem, shows how to construct spline discretisations that respect the budget, and establishes minimax lower bounds for operators constrained in both first and third derivative norms. The authors also demonstrate that layer errors can accumulate linearly with depth, indicating that the upper bound is not merely a theoretical artifact.

By Aleksander Tankman