The paper introduces SE(3) neural potential fields that learn collision‑free 6‑DoF trajectory planning directly from posed RGB images, eliminating the need for explicit 3D reconstruction. By supervising the field with a navigation function based on geodesic distances to the grasp, the method avoids the classic pitfalls of artificial potential fields, achieving near‑goal convergence within 3 cm from any start and producing collision‑free paths on a UR10 robot. Experiments on two tabletop scenes show significant improvements in clearance, reduced arm‑link contacts, and a 90 % grasp success rate, while planning time drops from over a minute to about 2 seconds compared to RRT* on a reconstructed scene.
By Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
arXiv:2606. 18634v1 Announce Type: cross Abstract: To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots.
By Zecheng Yin, Benedict Jun Ma
arXiv:2607. 18200v1 Announce Type: cross Abstract: Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts, while permissive margins lead to near-boundary shortcuts under perception bias.
By Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala, Anthony Tzes, Yi Fang
The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.
By Edward Beng Wai Tan, Siew-Kei Lam
arXiv:2609.22762v1 Announce Type: new
Abstract: Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expe...
By Fengcheng Yu, Dhruv Parikh, Junjie Ye, Maulik Bhatt, Thang Vu, Igor Vasiljevic, Vitor Guizilini, Yue Wang
Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction, and anticipatory planning need the future surface: the geometry at times beyond those captured. No standard benchmark measures this.