arXiv:2608.20948v1 Announce Type: cross
Abstract: Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this...
By Zhitao Liu, Guangtong Xu, Zihan Wang, Jialiang Hou, Chao Xu, Fei Gao
arXiv:2607. 20743v1 Announce Type: cross Abstract: Trajectory planning is a fundamental problem in robotics, requiring the generation of collision-free and efficient trajectories in a potentially complex environment.
By Miroslav Krupa, Miroslav Cibula, Krist\'ina Malinovsk\'a
The paper presents a training-free diffusion-based motion planner that replaces learned global trajectory scores with analytical local scores derived from obstacle, smoothness, velocity, and inter-agent feasibility terms. By reconstructing trajectory scores through local interactions between neighboring waypoints and nearby constraints, the method decomposes the denoising process while preserving the optimization structure of classical trajectory methods. Experiments demonstrate that this approach generates smooth, feasible trajectories for large multi-agent tasks in complex environments quickly, outperforming learning-based and optimization baselines without requiring training data.
By Michael Y. Fatemi, Jinhao Liang, Ferdinando Fioretto
The paper introduces Planning Diffusion Policy Optimization (PDPO), an offline‑to‑online reinforcement‑learning framework that employs a diffusion policy to produce short‑horizon action chunks for robot crowd navigation. PDPO is pretrained on collision‑avoidance demonstrations and fine‑tuned online with PPO, generating five‑step action sequences applied in a receding‑horizon manner. The authors also identify a benchmark artifact where agents can leave the valid domain without explicit boundary constraints, and they mitigate this by treating boundary violations as collisions, leading to improved success rates over strong baselines.
By Wendong Li, Jochen Garcke
arXiv:2607. 17574v1 Announce Type: cross Abstract: Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, with limited capacity to anticipate how moving obstacles will evolve in the near future.
By Yancheng Zhu, Wanli Ma, Chen Han, Irvin Haozhe Zhan, Bingfeng Qin, Yixin Xu
arXiv:2606. 01098v1 Announce Type: cross Abstract: Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control.
By Zemin Yang, Yaoyu He, Yiming Zhong, Yuhao Zhang, Xinge Zhu, Yao Mu, Qingqiu Huang, Yuexin Ma
arXiv:2604. 12474v3 Announce Type: replace-cross Abstract: In many robotic tasks, agents must traverse a sequence of spatial regions to complete a mission.
By Lidor Erez, Shahaf S. Shperberg, Ayal Taitler
arXiv:2602. 05031v2 Announce Type: replace Abstract: Planning with a learned model remains a key challenge in model-based reinforcement learning (RL).
By Dikshant Shehmar, Matthew Schlegel, Matthew E. Taylor, Marlos C. Machado
arXiv:2609.38383v1 Announce Type: cross
Abstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without po...
By Deqian Kong, Guangyan Sun, Sheng Cheng, Sirui Xie, Bo Pang, Jianwen Xie, Tony Geng, Caiwen Ding, Ying Nian Wu
arXiv:2603. 08862v2 Announce Type: replace-cross Abstract: Autonomous navigation in highly constrained environments remains challenging for mobile robots.
By Yuanjie Lu, Beichen Wang, Zhengqi Wu, Yang Li, Xiaomin Lin, Chengzhi Mao, Xuesu Xiao
arXiv:2606. 18828v1 Announce Type: cross Abstract: Traditional approaches place intelligence in the agent, whether as a learned policy or a search procedure.
By Chenghao Xu
Reinforced Planning with Latent World Models (RP1) is a novel method that learns to evaluate imagined outcomes via a critic and to improve multi‑step plans through an optimizer trained offline on world‑model roll‑outs. It is the first approach to fully learn plan improvement and can be attached to any pretrained latent world model. In experiments on visual navigation, arm reaching, and robotic manipulation, RP1 outperforms hand‑designed search algorithms, achieving near‑perfect success while using far fewer roll‑outs and running up to 67× faster than the strongest alternative.
By Armin Sommer, Jannik Schilling