The paper introduces a continual‑learning framework for single‑view 6‑DoF grasp synthesis with a parallel‑jaw gripper in cluttered scenes. Instead of fine‑tuning a large parametric model, the method updates grasp scores via memory in a learned embedding space and optionally incorporates user demonstrations to generate new candidate grasps. Experiments in simulation and real‑world trials (over 1500 grasps) show that the approach matches baseline performance before adaptation, improves online on unseen objects, and achieves over 90% success on challenging categories after just 50 online attempts.
By Giulio Schiavi, Andrei Cramariuc, Michael Pantic, Roland Siegwart
arXiv:2608.22449v1 Announce Type: cross
Abstract: Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based methods typical...
By Wen Wang, Ruibing Hou, Hong Chang, Shiguang Shan, Xilin Chen
arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.
By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin
arXiv:2609.24660v2 Announce Type: cross
Abstract: Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot...
By Shengcheng Luo, Xiaoyang Cheng, Hong Ying, Xiaoying Zhou, Jiaming Jiang, Haoran Guo, Wanlin Li, Ziyuan Jiao, Chenxi Xiao
arXiv:2608. 19759v1 Announce Type: cross Abstract: Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets.
By Julien Merand, Boris Meden, Mathieu Grossard, Liming Chen
arXiv:2606. 26428v1 Announce Type: cross Abstract: Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach.
By Tyler Ga Wei Lum, Kushal Kedia, C. Karen Liu, Jeannette Bohg
arXiv:2608. 15917v1 Announce Type: cross Abstract: Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers.
By Sarthak Kamat, Adam Rashid, Satvik Sharma, Aseem Doriwala, Chelsea Finn, Phillip Isola, C. Karen Liu
arXiv:2602. 13197v2 Announce Type: replace-cross Abstract: The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new source of highly scalable data for robot learning.
By Albert J. Zhai, Kuo-Hao Zeng, Jiasen Lu, Ali Farhadi, Shenlong Wang, Wei-Chiu Ma
arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.
By Jonghoon Lee, Seong Hyeon Park, Byungwoo Jeon, Minha Lee, Jinwoo Shin
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
FlowCorrect is a modular interactive imitation learning method that allows real‑time adaptation of generative flow‑matching manipulation policies using sparse, relative human corrections. During task execution, a human provides brief corrective pose nudges through a lightweight VR interface, and FlowCorrect locally adapts the policy without retraining the backbone, maintaining performance on previously learned scenarios. Experiments on a real‑world robot across four tabletop tasks show that, with a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on solved scenarios.
By Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.
By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu