arXiv Machine Learning By Yeongjong Kim, Minseok Kim, Yeoneung Kim, Namkyeong Cho

Physics-Informed Policy Iteration for High-Dimensional Hamilton--Jacobi--Bellman Equations: Interior Error Bounds without Boundary Data

Read the original on arXiv Machine Learning →

arXiv:2508. 01718v2 Announce Type: replace Abstract: We develop a physics-informed policy-iteration method for stationary second-order Hamilton--Jacobi--Bellman equations arising in continuous-time stochastic control.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration

The paper introduces a mesh‑free policy iteration framework that blends classical dynamic programming with physics‑informed neural networks (PINNs) to solve high‑dimensional, nonconvex Hamilton–Jacobi–Isaacs (HJI) equations. The method alternates between solving linear second‑order PDEs under fixed feedback policies and updating controls via pointwise minimax optimization using automatic differentiation. The authors prove local uniform convergence of the value function iterates to the unique viscosity solution under standard Lipschitz and uniform ellipticity assumptions, and demonstrate the approach’s accuracy and scalability in two‑, five‑, and ten‑dimensional stochastic games, outperforming direct PINN solvers.

By Hee Jun Yang, Minjung Gim, Yeoneung Kim
arXiv Machine Learning
Jun 25

A Zeroth-Order Deep Learning Method for Fully Nonlinear Parabolic Partial Differential Equations with Unknown Coefficients

arXiv:2606. 24999v1 Announce Type: new Abstract: High-dimensional partial differential equations (PDEs) with unknown coefficients arise widely in scientific machine learning, including continuous-time reinforcement learning, yet solving them efficiently in a data-driven way remains challenging.

By Yanwei Jia, Du Ouyang, Huy\^en Pham, Xun Yu Zhou
arXiv Machine Learning
Sep 17

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

The paper presents a convergence framework for deep $V$‑learning over a finite horizon $H$, deriving explicit bounds on policy loss by decomposing the Bellman update error into six residuals. It shows how $L^s$ concentrability controls expected $L^1$ loss, quantifies the impact of shared sampling across horizon levels, and provides optimal and near‑optimal sample allocations for statistical error rates. The work also establishes sharp action‑gap bounds under a margin condition, transfers optimal‑gap results to frozen‑iterate gaps, and offers consistency guarantees for generative‑reset approximate‑ERM procedures with exact action scores.

By Yury Kolomeytsev