arXiv AI

Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking

arXiv:2607. 17757v1 Announce Type: cross Abstract: Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in unstructured environments.

arXiv Machine Learning
Jul 7

Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance

arXiv:2603. 07866v3 Announce Type: replace-cross Abstract: Offshore inspection and maintenance have increasingly been using legged robots for routine sensing, yet many useful interventions still require physical interaction with tools, containers, and task-relevant objects.

By Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto, Ranulfo Bezerra, Gustavo J. G. Lahr, Ricardo V. Godoy, Marcelo Becker
arXiv AI
Jun 19

Human Universal Grasping

arXiv:2606. 17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality.

By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
arXiv Computer Vision
Sep 3

MultiGraspNet: A Multitask 3D Vision Model for Multi-gripper Robotic Grasping

MultiGraspNet is a multitask 3D vision model that simultaneously predicts feasible poses for both parallel and vacuum grippers, allowing a single robot to handle multiple end effectors. Trained on the aligned GraspNet-1Billion and SuctionNet-1Billion datasets, it generates graspability masks that quantify the suitability of each scene point for successful grasps. With only 15.75 M parameters, the model achieves fast inference on a single GPU and demonstrates competitive performance against single-task models while reducing computational cost, as shown in extensive experiments and real‑world tests on a single‑arm multi‑gripper setup.

By Stephany Ortuno-Chanelo, Paolo Rabino, Enrico Civitelli, Tatiana Tommasi, Raffaello Camoriano
arXiv Computer Vision
6d ago

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.

By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv AI
Jun 17

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

arXiv:2605. 31286v2 Announce Type: replace-cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments.

By Taiyi Su, Jian Zhu, Tianjian Wang, Youzhang He, Zitai Huang, Jianjun Zhang, Chong Ma, Hanyang Wang, Tianjiao Zhang, Munan Yin, Weihao Ding, Yi Xu
arXiv AI
Aug 19

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

The paper presents a perception-action pipeline for robotic pick‑and‑place in convenience stores, using annotation‑guided visual prompting to identify pickable objects and placement locations via bounding boxes. It replaces traditional step‑by‑step planning with Action Chunking with Transformers (ACT), an imitation learning algorithm that predicts chunked action sequences from human demonstrations. The system is evaluated on success rate and visual analysis of grasping behavior, showing improved grasp accuracy and adaptability in retail environments.

By Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae