arXiv AI

VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents

arXiv:2606. 19729v1 Announce Type: cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.

arXiv AI
Jun 16

PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty

arXiv:2606. 15654v1 Announce Type: cross Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive.

By Wenjing Tang, Xuanjin Jin, Yuan Liu, Renming Huang, Cewu Lu, Panpan Cai
arXiv Machine Learning
Aug 28

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

The paper introduces Planning Diffusion Policy Optimization (PDPO), an offline‑to‑online reinforcement‑learning framework that employs a diffusion policy to produce short‑horizon action chunks for robot crowd navigation. PDPO is pretrained on collision‑avoidance demonstrations and fine‑tuned online with PPO, generating five‑step action sequences applied in a receding‑horizon manner. The authors also identify a benchmark artifact where agents can leave the valid domain without explicit boundary constraints, and they mitigate this by treating boundary violations as collisions, leading to improved success rates over strong baselines.

By Wendong Li, Jochen Garcke
Hugging Face Trending Papers
Aug 20

RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation

Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion- and flow-matching robot policies lack tractable likelihoods, limiting their use in likelihood-based offline RL post-training.

arXiv Machine Learning
Aug 20

Reinforced Planning with Latent World Models

Reinforced Planning with Latent World Models (RP1) is a novel method that learns to evaluate imagined outcomes via a critic and to improve multi‑step plans through an optimizer trained offline on world‑model roll‑outs. It is the first approach to fully learn plan improvement and can be attached to any pretrained latent world model. In experiments on visual navigation, arm reaching, and robotic manipulation, RP1 outperforms hand‑designed search algorithms, achieving near‑perfect success while using far fewer roll‑outs and running up to 67× faster than the strongest alternative.

By Armin Sommer, Jannik Schilling
arXiv AI
Sep 1

Learning Action Models with Conditional and Quantified Effects via Uncertainty-Guided Exploration

The paper introduces OHCAM, an online method for learning action models that include conditional and quantified effects from limited interactions. It maintains a belief over possible models and actively chooses actions that maximize disagreement among hypotheses to reduce uncertainty, while handling noisy observations. Starting with simple hypotheses, OHCAM expands complexity only when necessary, achieving sample‑efficient learning that outperforms baselines on benchmark domains and is validated on a Kinova Gen3 robot.

By Jeffrey Jewett, William Solow, Sandhya Saisubramanian