The paper introduces SE(3) neural potential fields that learn collision‑free 6‑DoF trajectory planning directly from posed RGB images, eliminating the need for explicit 3D reconstruction. By supervising the field with a navigation function based on geodesic distances to the grasp, the method avoids the classic pitfalls of artificial potential fields, achieving near‑goal convergence within 3 cm from any start and producing collision‑free paths on a UR10 robot. Experiments on two tabletop scenes show significant improvements in clearance, reduced arm‑link contacts, and a 90 % grasp success rate, while planning time drops from over a minute to about 2 seconds compared to RRT* on a reconstructed scene.
By Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
arXiv:2609.31374v1 Announce Type: cross
Abstract: Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded...
By Zijun Zhao, Liewen Liao, Kang Shen, Songan Zhang, Ming Yang
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
arXiv:2610.01019v1 Announce Type: new
Abstract: Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative p...
By Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
arXiv:2606.21596v2 Announce Type: replace
Abstract: Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that...
By Haodong Li, Lulu Shao, Haolin Lu, Yu Fu, Yen-Ru Chen, Seemandhar Jain, Manmohan Chandraker
arXiv:2608. 11521v1 Announce Type: cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency.
By Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li