arXiv AI
Sep 4

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.

By Sixu Yan, Shikang Wang, Binhua Huang, Xuanlai Tang, Guohua Fan, Fan Huang, Haoxuan Li, Yongkang Li, Yuhan Li, Bencheng Liao, Zeyu Zhang, Wenyu Liu, Hangxin Liu, Xinggang Wang
arXiv Machine Learning
Sep 22

Connectivity-Aware Exploration of Robotic Grasp Spaces

The paper investigates the multiscale connectivity structure of successful robotic grasps in SE(3) and finds that these sets exhibit heterogeneous yet reproducible connectivity across objects. It proposes a connectivity‑aware sampling strategy that prioritizes bridges, frontiers, boundary extensions, and geometric novelty, which recovers grasp connectivity more efficiently than random or farthest‑point sampling. Experiments also show that connectivity information can improve subsequent grasp discovery and transfer to unseen objects, indicating that the spatial organization of viable actions offers valuable guidance for exploration.

By Maksim A Kazanskii