arXiv:2606. 15271v1 Announce Type: cross Abstract: This work presents a transparent and reproducible benchmark study of a direct dual-network Physics-Informed Neural Network (PINN) formulation for the optimal control of a mass-spring-damper system.
By Abdeladhim Tahimi, Rinaldo Vieira da Silva Junior
The paper presents a convergence framework for deep $V$‑learning over a finite horizon $H$, deriving explicit bounds on policy loss by decomposing the Bellman update error into six residuals. It shows how $L^s$ concentrability controls expected $L^1$ loss, quantifies the impact of shared sampling across horizon levels, and provides optimal and near‑optimal sample allocations for statistical error rates. The work also establishes sharp action‑gap bounds under a margin condition, transfers optimal‑gap results to frozen‑iterate gaps, and offers consistency guarantees for generative‑reset approximate‑ERM procedures with exact action scores.
By Yury Kolomeytsev
arXiv:2609.24489v1 Announce Type: new
Abstract: Offline reinforcement learning (RL) typically trains a critic by minimizing a regression loss against bootstrapped value targets stabilized by target n...
By Hyukjun Yang, Jongchan Park, Narim Jeong, Donghwan Lee
arXiv:2606. 25680v2 Announce Type: replace-cross Abstract: Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task while drawing less thruster power directly extends mission range and endurance.
By Yinuo Wang, Gavin Tao, Yuze Liu, John V. Ringwood
arXiv:2605. 26790v3 Announce Type: replace Abstract: Low-thrust trajectory design relies heavily on repeated evaluations of fuel consumption and transfer feasibility, which require expensive optimal control solutions.
By Zhong Zhang, Giacomo Acciarini, Dario Izzo, Hexi Baoyin, Francesco Topputo
arXiv:2603. 18853v3 Announce Type: replace-cross Abstract: Autonomous aerial vehicles (AAVs) enable data collection for sixth-generation Internet-of-Things networks, but their trajectories couple nonlinear wireless rates with long-horizon service progress.
By Xiucheng Wang, Zhenye Chen, Nan Cheng, Zhisheng Yin, Xuemin Shen