arXiv AI

GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning

arXiv:2509. 11594v3 Announce Type: replace-cross Abstract: GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapshot.

arXiv Machine Learning
1d ago

Continual Learning for 6-DoF Grasp Synthesis via Experience and Demonstrations

The paper introduces a continual‑learning framework for single‑view 6‑DoF grasp synthesis with a parallel‑jaw gripper in cluttered scenes. Instead of fine‑tuning a large parametric model, the method updates grasp scores via memory in a learned embedding space and optionally incorporates user demonstrations to generate new candidate grasps. Experiments in simulation and real‑world trials (over 1500 grasps) show that the approach matches baseline performance before adaptation, improves online on unseen objects, and achieves over 90% success on challenging categories after just 50 online attempts.

By Giulio Schiavi, Andrei Cramariuc, Michael Pantic, Roland Siegwart
arXiv AI
Jun 19

Human Universal Grasping

arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.

By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv AI
Sep 4

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.

By Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
arXiv Computer Vision
6d ago

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.

By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv AI
Jun 9

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.

By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu
Hugging Face Trending Papers
5d ago

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

The paper introduces P2P‑T, a data‑efficient, object‑centric framework that learns tool manipulation directly from human video demonstrations. It uses a two‑stage approach: first pretraining an object‑centric world model to extract stable pose priors, then integrating these priors into a pose‑aware low‑level policy. By automating data processing with foundation models, P2P‑T eliminates the need for human‑robot aligned data and achieves a 73% improvement over prior state‑of‑the‑art performance on complex real‑world tool manipulation tasks.

arXiv AI
Sep 21

2nd Place Solution to the HANDS 2026 Workshop Challenge-Dexterous Grasp Motion Track: Single-Shot Trajectory Warping for Grasp Motion Generation

The paper reports a 2nd place solution for the HANDS 2026 Dexterous Grasp Motion track, targeting the 12‑DoF LinkerHand O6. The method warps a single successful GraspM3 demonstration into a 12‑D trajectory, avoiding step‑by‑step policy generation, and is trained with one‑step PPO across 4,824 objects. It achieved 94.61% success on the easy track and 57.18% on the hard track of the private test set.

By Muneeb A. Khan, Woojin Kim, Shinwoo Kim, Muhammad Munsif, Binod Bhattarai, Seungryul Baek