The paper introduces a framework that uses world models to enable robotic insertion across diverse parts. By combining proprioceptive data with wrist‑mounted camera visuals, a single model is trained on up to 90 tasks, achieving 56% zero‑shot success on unseen objects versus 7% for a model‑free baseline. The approach scales with more training objects and can be fine‑tuned for improved data efficiency and performance.
By Nicklas Hansen, Iretiayo Akinola, Yijie Guo, Jie Xu, Bingjie Tang, Hao Su, Xiaolong Wang, Abhishek Gupta, Dieter Fox, Yashraj Narang
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv:2604. 10579v2 Announce Type: replace-cross Abstract: Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations due to limited data diversity.
By Jiawei Zhang, Kaizhe Hu, Yingqian Huang, Yuanchen Ju, Zhengrong Xue, Huazhe Xu
arXiv:2509. 06191v2 Announce Type: replace-cross Abstract: Recent 3D generative models, which are capable of generating full object shapes from just a few images, now open up new opportunities in robotics.
By Yifei Ren, Edward Johns
Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.
KeyGen is a framework that learns canonical 3D keypoints from point clouds to create structured, object‑centric representations for policy learning in robotic manipulation. By conditioning a visuomotor diffusion policy on these keypoints and object geometry, it predicts full manipulation trajectories that maintain geometric correspondence across different object instances. Experiments on a photorealistic simulation benchmark with three tasks show that KeyGen outperforms prior methods on both seen and unseen objects, scales with more demonstrations, remains robust to rescaling, and performs well in real‑world manipulation.
By Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg
arXiv:2608. 19968v1 Announce Type: cross Abstract: Modern computer vision has enabled partial autonomy in robotic assembly manipulation.
By Kulunu Samarawickrama, Roel Pieters
arXiv:2606. 08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments.
By Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan
arXiv:2608. 19776v1 Announce Type: cross Abstract: Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks.
By Julien Merand, Boris Meden, Liming Chen, Mathieu Grossard
arXiv:2609.38443v1 Announce Type: cross
Abstract: We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yi...
By Cameron Smith, Arsh Tangri, Vitor Guizilini, Yue Wang, Zubair Irshad, Sergey Zakharov
arXiv:2610.01758v1 Announce Type: new
Abstract: Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene...
By Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao
arXiv:2608. 19759v1 Announce Type: cross Abstract: Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets.
By Julien Merand, Boris Meden, Mathieu Grossard, Liming Chen