arXiv Computer Vision

Vision-Based 6-DoF Grasp Pose Estimation for Robot Cloth Unfolding

arXiv AI
Jun 17

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

arXiv:2605. 31286v2 Announce Type: replace-cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments.

By Taiyi Su, Jian Zhu, Tianjian Wang, Youzhang He, Zitai Huang, Jianjun Zhang, Chong Ma, Hanyang Wang, Tianjiao Zhang, Munan Yin, Weihao Ding, Yi Xu
arXiv Computer Vision
Sep 3

MultiGraspNet: A Multitask 3D Vision Model for Multi-gripper Robotic Grasping

MultiGraspNet is a multitask 3D vision model that simultaneously predicts feasible poses for both parallel and vacuum grippers, allowing a single robot to handle multiple end effectors. Trained on the aligned GraspNet-1Billion and SuctionNet-1Billion datasets, it generates graspability masks that quantify the suitability of each scene point for successful grasps. With only 15.75 M parameters, the model achieves fast inference on a single GPU and demonstrates competitive performance against single-task models while reducing computational cost, as shown in extensive experiments and real‑world tests on a single‑arm multi‑gripper setup.

By Stephany Ortuno-Chanelo, Paolo Rabino, Enrico Civitelli, Tatiana Tommasi, Raffaello Camoriano
arXiv AI
Jun 9

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

arXiv:2606. 08057v1 Announce Type: cross Abstract: Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often missing or require pre-scanned object assets.

By Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan, Shiyu Gao, Ying Ai, Hui Xu, Yongqi Hu, Hengyi Zhang, Yang Xie, Zhaxizhuoma, Yue Zhao, Zhenshan Bing, Yan Ding, Jianxing Liu
arXiv Computer Vision
Sep 28

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

The paper introduces an interaction‑centric framework that unifies representations for two‑finger gripper manipulation across different robot embodiments. By using a parameterized universal gripper abstraction and a canonical gripper‑frame representation, the system infers sub‑tasks from language and RGB‑D inputs, grounds interaction triplets, and employs hybrid features and a Flow‑Matching Transformer to generate smooth 7‑DoF action sequences. Experiments in both simulation and real‑world settings show that this approach achieves competitive benchmark performance while enabling extreme cross‑embodiment and cross‑viewpoint zero‑shot sim‑to‑real transfer to heterogeneous robot platforms.

By Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
arXiv AI
Jul 2

Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration

arXiv:2607. 00033v1 Announce Type: cross Abstract: Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging.

By Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Milad Noori, Huihua Zhao, John Welsh, Michael Andres Lin, Wei Liu, Tingwu Wang, Xingye Da, Zhengyi Luo, Vishal Kulkarni, Naema Bhatti, Yuke Zhu, Linxi Fan, Bowen Wen, Danfei Xu, Soha Pouya, Yan Chang
arXiv AI
Jul 1

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

arXiv:2606. 31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse.

By Tao Chen, Lizheng Liu, Jiaxu Wang, Ziyue Jiang, Ruiqi Tian, JiGuang Huo, Zhongxue Gan