arXiv Computer Vision

TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

arXiv Computer Vision
1d ago

UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision

arXiv:2609.39388v1 Announce Type: cross Abstract: Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in acti...

By Wei Xue, Keliang Liu, Mingzhang Cui, Jinhua Xie, Jinjie Wei, Jianan Hou, Jingcheng Lu, Lintao Wang, Kaixiang Qiu, Yizhou Liu, Xinghai Ye, Jinghang Han, Mingcheng Li, Jie Gu, Shunli Wang, Lihua Zhang, Dingkang Yang
Hugging Face Trending Papers
Jul 6

Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, language understanding, and robot action generation. Existing research has primarily focused on improving model architectures, training strategies, and dataset scale, while little attention has been paid to how demonstrations are collected and organized.

arXiv AI
Jul 2

Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration

arXiv:2607. 00033v1 Announce Type: cross Abstract: Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging.

By Xinghao Zhu, Zixi Liu, Shalin Jain, Chenran Li, Milad Noori, Huihua Zhao, John Welsh, Michael Andres Lin, Wei Liu, Tingwu Wang, Xingye Da, Zhengyi Luo, Vishal Kulkarni, Naema Bhatti, Yuke Zhu, Linxi Fan, Bowen Wen, Danfei Xu, Soha Pouya, Yan Chang
arXiv AI
Jun 9

GEAR-VLA: Learning Geometry-Aware Action Representations for Generalizable Robotic Manipulation

arXiv:2606. 08530v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong benchmark performance but still struggle in real-world deployment with unseen objects, background shifts, and different robot embodiments.

By Yuan Zhang, Shiqi Zhang, Yedong Shen, Shuai Dong, Jiajun Deng, Xin Zhang, Yuxuan Gao, Jiajia Wu, Xin Nie, Zhiyuan Cheng, Jianmin Ji, Yanyong Zhang, Xingyi Zhang, Jia Pan
arXiv Machine Learning
Sep 23

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

MAVP (Map-Aware Visuomotor Policies) is a framework that enhances mobile manipulation by predicting explicit base-pose targets and tracking them with localisation feedback. It reconstructs a static map from teleoperated demonstrations, expressing base trajectories in a shared map frame to provide consistent spatial supervision. During execution, the policy uses RGB observations, joint states, and the robot’s current map-frame base pose to jointly predict target base poses, arm actions, and gripper actions, while a low‑level controller corrects deviations using feedforward motion and pose error feedback. Pose‑noise augmentation during training further improves robustness, and MAVP outperforms unanchored velocity control across six real‑world tasks and three policy families.

By Jinhe Tang, Ruixiao Dai, Weiming Zhi