arXiv Machine Learning

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

MAVP (Map-Aware Visuomotor Policies) is a framework that enhances mobile manipulation by predicting explicit base-pose targets and tracking them with localisation feedback. It reconstructs a static map from teleoperated demonstrations, expressing base trajectories in a shared map frame to provide consistent spatial supervision. During execution, the policy uses RGB observations, joint states, and the robot’s current map-frame base pose to jointly predict target base poses, arm actions, and gripper actions, while a low‑level controller corrects deviations using feedforward motion and pose error feedback. Pose‑noise augmentation during training further improves robustness, and MAVP outperforms unanchored velocity control across six real‑world tasks and three policy families.

arXiv Computer Vision
1d ago

UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision

arXiv:2609.39388v1 Announce Type: cross Abstract: Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in acti...

By Wei Xue, Keliang Liu, Mingzhang Cui, Jinhua Xie, Jinjie Wei, Jianan Hou, Jingcheng Lu, Lintao Wang, Kaixiang Qiu, Yizhou Liu, Xinghai Ye, Jinghang Han, Mingcheng Li, Jie Gu, Shunli Wang, Lihua Zhang, Dingkang Yang
arXiv Machine Learning
Aug 31

FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation

FlowCorrect is a modular interactive imitation learning method that allows real‑time adaptation of generative flow‑matching manipulation policies using sparse, relative human corrections. During task execution, a human provides brief corrective pose nudges through a lightweight VR interface, and FlowCorrect locally adapts the policy without retraining the backbone, maintaining performance on previously learned scenarios. Experiments on a real‑world robot across four tabletop tasks show that, with a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on solved scenarios.

By Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
arXiv AI
Sep 18

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.

By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
arXiv Computer Vision
Aug 25

TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

arXiv:2608.22296v1 Announce Type: cross Abstract: Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact th...

By Haoran Lin, Mingyu Yang, Pengfei Qi, Kehan Chen, Qiang Diao, Liangji Zeng, Wenrui Chen, Yaonan Wang, Kailun Yang
Hugging Face Trending Papers
Jul 27

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations.

arXiv Machine Learning
Jun 10

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations

arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.

By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin