arXiv Machine Learning By Minseok Kim, Yeongjong Kim, Namkyeong Cho, Yeoneung Kim

Monotone Neural Policy Iteration for High-Dimensional First-Order Hamilton--Jacobi--Bellman Equations

Read the original on arXiv Machine Learning →

arXiv:2605. 07116v2 Announce Type: replace Abstract: We analyze a neural semi-discrete method for high-dimensional first-order Hamilton-Jacobi-Bellman (HJB) equations with known or learned dynamics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration

The paper introduces a mesh‑free policy iteration framework that blends classical dynamic programming with physics‑informed neural networks (PINNs) to solve high‑dimensional, nonconvex Hamilton–Jacobi–Isaacs (HJI) equations. The method alternates between solving linear second‑order PDEs under fixed feedback policies and updating controls via pointwise minimax optimization using automatic differentiation. The authors prove local uniform convergence of the value function iterates to the unique viscosity solution under standard Lipschitz and uniform ellipticity assumptions, and demonstrate the approach’s accuracy and scalability in two‑, five‑, and ten‑dimensional stochastic games, outperforming direct PINN solvers.

By Hee Jun Yang, Minjung Gim, Yeoneung Kim
arXiv Machine Learning
Jun 2

A Per-Component Diagnostic Protocol for Neural HJB-PIDE Solvers under Control-Dependent L\'evy Jumps

arXiv:2606. 01122v1 Announce Type: new Abstract: We propose a five-step diagnostic protocol for residual-trained neural HJB-PIDE solvers with control-dependent L\'evy jumps, targeting a general failure mode of neural PDE methods: a learned solution can match headline scalar diagnostics while miscomputing an operator inside its training loss.

By R. Drissi
arXiv Machine Learning
Jun 26

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

arXiv:2606. 26498v1 Announce Type: cross Abstract: This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available.

By Erhan Bayraktar, Martin Hernandez, Qinxin Yan, Yuhua Zhu
Hugging Face Trending Papers
Jun 25

Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based formulation, policy evaluation is naturally described by a stationary Hamilton-Jacobi-Bellman equation on $\mathcal P_2(\mathbb R^d)$, but this equation involves the drift and diffusion coefficients of the controlled McKean-Vlasov dynamics, which are not identifiable when only discrete-time data are available.