arXiv:2607. 13553v1 Announce Type: cross Abstract: Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments.
By Andrea Maria Braghin, Nicol\`o Botteghi, Matteo Tomasetto, Andrea Manzoni, Gabriele Cazzulani
arXiv:2607. 14272v1 Announce Type: new Abstract: Flow matching has emerged as an effective framework for learning complex data distributions, but adapting pretrained flow models to new tasks often requires computationally expensive retraining.
By Jingdong Zhang, Xinze Li, Yize Jiang, Luan Yang, Minkai Xu, Junhong Liu
The paper introduces Multi-step Proximal Policy Improvement (MPI), a method that refines offline reinforcement learning policies through sequential re-centered proximal steps. By viewing policies as a probability manifold, MPI interprets a wide range of offline actor objectives as a single proximal policy improvement step and extends this to multiple steps for controlled policy improvement beyond the behavior distribution. Experiments on D4RL benchmarks demonstrate that a few MPI refinements enhance strong offline baselines such as TD3+BC, ReBRAC, and IQL, while diagnostics clarify the benefits of re-centered refinement over fixed-objective scheduling and highlight critic error limitations.
By Soohyun Choi, Seonvin Cho, Songnam Hong
The paper introduces Multi-step Proximal Policy Improvement (MPI), a method that refines offline reinforcement learning policies through sequential re-centered proximal steps. By modeling policies as a probability manifold, MPI interprets a wide range of offline actor objectives as a single proximal policy improvement step and extends this to multiple steps for controlled policy improvement beyond the behavior distribution. Experiments on D4RL benchmarks demonstrate that a few MPI refinements enhance strong offline baselines such as TD3+BC, ReBRAC, and IQL across many tasks, while diagnostics clarify the benefits of re-centered refinement over fixed-objective scheduling and highlight critic error limitations.
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2512. 17534v2 Announce Type: replace-cross Abstract: Modeling and controlling fluids is critical across science and engineering.
By Christian Lagemann, Sajeda Mokbel, Miro Gondrum, Mario R\"uttgers, Yuning Wang, Pol Su\'arez, Ludger Paehler, Deniz A. Bezgin, Aaron B. Buhendwa, Jared L. Callaham, Samuel Ahnert, Nicholas Zolman, Xiao Shao, Jean-Christophe Loiseau, Nikolaus Adams, Matthias Meinke, Wolfgang Schr\"oder, Kai Lagemann, Esther Lagemann, Ricardo Vinuesa, Steven L. Brunton
arXiv:2107. 08183v2 Announce Type: replace Abstract: High-dimensional state and action spaces combined with sparse reward structures in reinforcement learning (RL) environments typically require advanced control architectures.
By JaeYoon Kim, Junyu Xuan, Christy Liang, Farookh Hussain
arXiv:2607. 16177v1 Announce Type: new Abstract: Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems.
By Matteo Tomasetto, Nicol\`o Botteghi, Gabriele Bruni, Andrea Manzoni
arXiv:2607. 21302v1 Announce Type: new Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations.
By Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia
arXiv:2606. 29820v1 Announce Type: cross Abstract: In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimation and multimodal exploration challenging.
By Qijun Li, Zheng Fu, Qi Song, Yifei He, Weitao Zhou, Kun Jiang, Diange Yang
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2609.15123v1 Announce Type: cross
Abstract: Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from...
By Bumgeun Park, Hyukjun Yang, Donghwan Lee