The paper introduces SE(3) neural potential fields that learn collision‑free 6‑DoF trajectory planning directly from posed RGB images, eliminating the need for explicit 3D reconstruction. By supervising the field with a navigation function based on geodesic distances to the grasp, the method avoids the classic pitfalls of artificial potential fields, achieving near‑goal convergence within 3 cm from any start and producing collision‑free paths on a UR10 robot. Experiments on two tabletop scenes show significant improvements in clearance, reduced arm‑link contacts, and a 90 % grasp success rate, while planning time drops from over a minute to about 2 seconds compared to RRT* on a reconstructed scene.
By Jeffrey Eiyike, Masoud Ataei, Elvis Gyaase, Vikas Dhiman
arXiv:2609.31374v1 Announce Type: cross
Abstract: Closed-loop driving simulation requires rendered observations to remain reliable as the ego vehicle and surrounding actors move beyond their recorded...
By Zijun Zhao, Liewen Liao, Kang Shen, Songan Zhang, Ming Yang
RoboPhys-3D is a 3D‑grounded embodied world model benchmark built on RoboTwin 2.0, featuring 50 manipulation tasks, 5,000 episodes, and 25,000 multi‑view ground‑truth videos. It evaluates video world models by processing both generated and ground‑truth videos through the same 3D reconstruction pipeline, allowing the separation of reconstruction‑induced from generation‑induced errors. The benchmark defines 50 metrics across four sub‑dimensions—pixel fidelity, 3D geometry consistency, state understanding, and task completeness—and introduces the Average Full Score and RoboPhyscore for holistic assessment, with RoboPhyscore showing strong correlation with human judgments.
By Tianyi Wang, Jiazhou Chen, Yiming Xu, Xiangyu Li, Tianyi Zeng, Chih-Hsien Chou, Ning Lu, Liang Peng, Junfeng Jiao, Christian Claudel
arXiv:2610.01019v1 Announce Type: new
Abstract: Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative p...
By Hao Wu, Shengju Qian, Weiyan Wang, Fan Xu, Fan Zhang, Yuanpeng He, Qingsong Wen, Yuxuan Liang
arXiv:2606.21596v2 Announce Type: replace
Abstract: Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that...
By Haodong Li, Lulu Shao, Haolin Lu, Yu Fu, Yen-Ru Chen, Seemandhar Jain, Manmohan Chandraker
arXiv:2608. 11521v1 Announce Type: cross Abstract: World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency.
By Chushan Zhang, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
arXiv:2606. 13053v1 Announce Type: cross Abstract: Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant events.
By Kailin Wang, Haoxiang Jie, Yaoyuan Yan, Jiacheng Zhou, Zhiyou Heng
Reaching a 6-DoF grasp pose in clutter requires a collision-free trajectory, conventionally obtained by reconstructing the scene in 3D and planning inside that reconstruction, at the cost of its accur...
arXiv:2608.31002v1 Announce Type: cross
Abstract: Robotic perception from a single viewpoint is often limited by self-occlusion and incomplete surface visibility. This paper presents DARP(Dual-Arm Ro...
By Manish Kansana, Mohammed Yusuf Mujawar, Sudip Mittal, Shahram Rahimi, Noorbakhsh Amiri Golilarz
arXiv:2605. 10873v2 Announce Type: replace-cross Abstract: Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure because existing evaluations are fragmented across datasets, modalities, and metrics.
By Anna C. Doris, Jacob Thomas Sony, Ghadi Nehme, Era Syla, Amin Heyrani Nobari, Faez Ahmed
arXiv:2605. 21862v2 Announce Type: replace-cross Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone.
By Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux