arXiv:2606. 25012v1 Announce Type: new Abstract: Many reinforcement learning (RL) problems in the infinite-horizon average-reward setting require optimizing multiple conflicting objectives while satisfying multiple safety constraints.
By Ankur Naskar, Swetha Ganesh, Vaneet Aggarwal
arXiv:2606. 15271v1 Announce Type: cross Abstract: This work presents a transparent and reproducible benchmark study of a direct dual-network Physics-Informed Neural Network (PINN) formulation for the optimal control of a mass-spring-damper system.
By Abdeladhim Tahimi, Rinaldo Vieira da Silva Junior
The paper introduces Policy Gradient Penalty (PGP), a single‑loop policy‑space method that enforces convex occupancy‑measure constraints via quadratic‑penalty regularization. PGP constructs pseudo‑rewards to estimate gradients of the penalized objective and uses the classical Policy Gradient Theorem, establishing smoothness and global last‑iterate convergence guarantees for an ε‑optimal constrained entropy value with ε‑bounded constraint violation. The authors validate PGP with ablations on a grid‑world benchmark and demonstrate scalability on two challenging continuous‑control tasks.
By Florian Wolf, Ilyas Fatkhullin, Niao He
arXiv:2605. 07116v2 Announce Type: replace Abstract: We analyze a neural semi-discrete method for high-dimensional first-order Hamilton-Jacobi-Bellman (HJB) equations with known or learned dynamics.
By Minseok Kim, Yeongjong Kim, Namkyeong Cho, Yeoneung Kim
arXiv:2607. 11122v1 Announce Type: cross Abstract: Implicit neural controllers (INCs) are static feedback laws that are evaluated through an algebraic fixed point {equation}; they include as special cases neural network controllers.
By Giuseppe C. Calafiore, Laurent El Ghaoui
arXiv:2506. 01226v3 Announce Type: replace-cross Abstract: We study parameterizations of stabilizing nonlinear policies for learning-based control.
By Nicholas H. Barbara, Ruigang Wang, Alexandre Megretski, Ian R. Manchester