Learning Motion Feasibility from Point Clouds in Cluttered Environments
arXiv:2606. 26700v1 Announce Type: cross Abstract: Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation.
arXiv:2509. 11594v3 Announce Type: replace-cross Abstract: GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapshot.
arXiv:2606. 26700v1 Announce Type: cross Abstract: Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation.
The paper introduces a continual‑learning framework for single‑view 6‑DoF grasp synthesis with a parallel‑jaw gripper in cluttered scenes. Instead of fine‑tuning a large parametric model, the method updates grasp scores via memory in a learned embedding space and optionally incorporates user demonstrations to generate new candidate grasps. Experiments in simulation and real‑world trials (over 1500 grasps) show that the approach matches baseline performance before adaptation, improves online on unseen objects, and achieves over 90% success on challenging categories after just 50 online attempts.
arXiv:2608. 19759v1 Announce Type: cross Abstract: Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets.
arXiv:2608. 19776v1 Announce Type: cross Abstract: Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks.
arXiv:2607. 14341v1 Announce Type: cross Abstract: Robust robotic grasping remains a fundamental challenge for complex real-world applications.
arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.
AdaRoboVLG is a Vision‑Language‑Grasp framework that separates a generalizable base grasp policy from task‑specific understanding. The base policy generates and evaluates physically feasible grasp candidates using kinematic mapping and force‑closure stability, while foundation‑model modules supply composable spatial, cognitive, and temporal priors that adapt grasp synthesis to different robotic hands and environments without retraining. Experiments show strong cross‑hand generalization, effective handling of diverse grasping challenges, and functional grasping in cluttered, dynamic settings.
The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.
arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.
The paper introduces P2P‑T, a data‑efficient, object‑centric framework that learns tool manipulation directly from human video demonstrations. It uses a two‑stage approach: first pretraining an object‑centric world model to extract stable pose priors, then integrating these priors into a pose‑aware low‑level policy. By automating data processing with foundation models, P2P‑T eliminates the need for human‑robot aligned data and achieves a 73% improvement over prior state‑of‑the‑art performance on complex real‑world tool manipulation tasks.
The paper reports a 2nd place solution for the HANDS 2026 Dexterous Grasp Motion track, targeting the 12‑DoF LinkerHand O6. The method warps a single successful GraspM3 demonstration into a 12‑D trajectory, avoiding step‑by‑step policy generation, and is trained with one‑step PPO across 4,824 objects. It achieved 94.61% success on the easy track and 57.18% on the hard track of the private test set.
arXiv:2609.39375v1 Announce Type: cross Abstract: A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests...