arXiv:2605. 07116v2 Announce Type: replace Abstract: We analyze a neural semi-discrete method for high-dimensional first-order Hamilton-Jacobi-Bellman (HJB) equations with known or learned dynamics.
By Minseok Kim, Yeongjong Kim, Namkyeong Cho, Yeoneung Kim
arXiv:2608. 11480v1 Announce Type: cross Abstract: Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is bottlenecked by the computational complexity of solving Hamilton-Jacobi-Isaacs variational inequality PDEs in high dimensions.
By Sungje Park, Stephen Tu
Hamilton-Jacobi (HJ) reachability provides a mathematically rigorous framework for safe control of dynamical systems, but its practical application is bottlenecked by the computational complexity of solving Hamilton-Jacobi-Isaacs variational inequality PDEs in high dimensions. Physics-informed neural networks (PINNs) have recently emerged as a promising alternative to classical mesh-based solvers, yet their performance is highly sensitive to the choice of collocation sampling.
arXiv:2508. 01718v2 Announce Type: replace Abstract: We develop a physics-informed policy-iteration method for stationary second-order Hamilton--Jacobi--Bellman equations arising in continuous-time stochastic control.
By Yeongjong Kim, Minseok Kim, Yeoneung Kim, Namkyeong Cho
arXiv:2607. 16177v1 Announce Type: new Abstract: Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems.
By Matteo Tomasetto, Nicol\`o Botteghi, Gabriele Bruni, Andrea Manzoni
arXiv:2606. 28671v1 Announce Type: new Abstract: Stackelberg differential games (SDGs) provide a powerful framework for hierarchical decision-making in stochastic and continuous-time environments, yet their solution remains computationally challenging due to the complexity of traditional dynamic programming and Hamilton-Jacobi-Bellman-Isaacs (HJBI) methods, especially in high-dimensional systems.
By Congde Hu, Danping Li, Lin Xu, Wenying Xu
arXiv:2509. 23960v2 Announce Type: replace-cross Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge.
By Manan Tayal, Aditya Singh, Shishir Kolathaya, Somil Bansal
arXiv:2606. 20442v1 Announce Type: new Abstract: Physics-Informed Neural Networks (PINNs) solve Partial Differential Equations (PDEs) by embedding physical laws into neural network training.
By Fedor Buzaev (HSE University), Dmitry Efremenko (HSE University), Egor Bugaev (HSE University), Andrei Ermakov (HSE University, AXXX), Denis Derkach (HSE University), Daria Pugacheva (HSE University, AXXX), Fedor Ratnikov (HSE University)
arXiv:2601. 04120v2 Announce Type: replace-cross Abstract: Optimal control of obstacle problems arises in a wide range of applications and is computationally challenging due to its nonsmoothness, nonlinearity, and bilevel structure.
By Yongcun Song, Shangzhi Zeng, Jin Zhang, Lvgang Zhang
arXiv:2310. 07211v2 Announce Type: replace Abstract: Regularization is a cornerstone of modern reinforcement learning.
By Zeyang Li, Chuxiong Hu, Yunan Wang, Guojian Zhan, Jie Li, Yao Lyu, Shengbo Eben Li
The paper proves that the deep Galerkin method (DGM) converges when applied to Hamilton‑Jacobi‑Bellman equations derived from finite‑state mean field control problems. By showing that the DGM loss can be driven arbitrarily low under sufficient regularity of the value function, and that a vanishing loss forces uniform convergence of the neural network approximators to the true value function on the simplex, the authors establish both existence and convergence results for the DGM. Numerical experiments further illustrate the method’s ability to handle high‑dimensional HJB equations.
By William Hofgard, Jingruo Sun, Asaf Cohen
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou