Domain randomization and generative models for robotic grasping
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2608. 19759v1 Announce Type: cross Abstract: Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets.
AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.
The paper investigates the multiscale connectivity structure of successful robotic grasps in SE(3) and finds that these sets exhibit heterogeneous yet reproducible connectivity across objects. It proposes a connectivity‑aware sampling strategy that prioritizes bridges, frontiers, boundary extensions, and geometric novelty, which recovers grasp connectivity more efficiently than random or farthest‑point sampling. Experiments also show that connectivity information can improve subsequent grasp discovery and transfer to unseen objects, indicating that the spatial organization of viable actions offers valuable guidance for exploration.
arXiv:2606. 26428v1 Announce Type: cross Abstract: Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach.
arXiv:2608. 19776v1 Announce Type: cross Abstract: Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks.