VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents
arXiv:2606. 19729v1 Announce Type: cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.
arXiv:2510. 27191v5 Announce Type: replace-cross Abstract: Planning under partial observability is an essential capability of autonomous robots.
arXiv:2606. 19729v1 Announce Type: cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.
arXiv:2606. 19729v2 Announce Type: replace-cross Abstract: Planning under uncertainty is an essential capability for autonomous robots.
arXiv:2609.35874v1 Announce Type: new Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost state...
arXiv:2609.01351v1 Announce Type: cross Abstract: Online planning under uncertainty remains a fundamental challenge for robotic systems operating in partially observable environments with high-dimens...
arXiv:2608. 06702v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets under strict real-time constraints.
arXiv:2609.39559v1 Announce Type: new Abstract: In this work we study the problem of MAPFC, a post-optimization step for Multi-Agent Path Finding (MAPF) plans where we are given a feasible plan produ...
arXiv:2609.37432v1 Announce Type: new Abstract: Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also sup...
arXiv:2603.08814v2 Announce Type: replace-cross Abstract: Long-horizon task planning for heterogeneous multi-robot systems is essential for deploying collaborative teams in real-world environments; y...
arXiv:2606. 15654v1 Announce Type: cross Abstract: Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive.
arXiv:2406. 09953v4 Announce Type: replace-cross Abstract: Dual-arm robots promise greater efficiency but require planning for complex tasks with nonlinear sub-task dependencies.
arXiv:2608.28995v1 Announce Type: cross Abstract: World models let robots imagine possible futures, but exploiting this capability for real-time control is bottlenecked by a representation misalignme...
Reinforced Planning with Latent World Models (RP1) is a novel method that learns to evaluate imagined outcomes via a critic and to improve multi‑step plans through an optimizer trained offline on world‑model roll‑outs. It is the first approach to fully learn plan improvement and can be attached to any pretrained latent world model. In experiments on visual navigation, arm reaching, and robotic manipulation, RP1 outperforms hand‑designed search algorithms, achieving near‑perfect success while using far fewer roll‑outs and running up to 67× faster than the strongest alternative.