arXiv:2609.39388v1 Announce Type: cross
Abstract: Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in acti...
By Wei Xue, Keliang Liu, Mingzhang Cui, Jinhua Xie, Jinjie Wei, Jianan Hou, Jingcheng Lu, Lintao Wang, Kaixiang Qiu, Yizhou Liu, Xinghai Ye, Jinghang Han, Mingcheng Li, Jie Gu, Shunli Wang, Lihua Zhang, Dingkang Yang
Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However, their predictions for fine-grained spatial and ro...
arXiv:2609.10506v1 Announce Type: cross
Abstract: Action-conditioned latent world models predict future visual representations, enabling zero-shot goal-conditioned robot planning and control. However...
By Nisarga Nilavadi, Ralf R\"omer, Moritz Reuss, Michael Krawez, Tobias J\"ulg, Angela P. Schoellig, Rudolf Lioutikov, Wolfram Burgard
arXiv:2606. 10025v1 Announce Type: cross Abstract: We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution.
By Sriram Krishna, Ben Eisner, Haotian Zhan, Ying Yuan, Haoyu Zhen, Chuang Gan, Shubham Tulsiani, David Held
arXiv:2608. 20114v1 Announce Type: new Abstract: Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control.
By Siyuan Ma, Boshi Zhang, Yutian Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Qiaojun Yu
FlowCorrect is a modular interactive imitation learning method that allows real‑time adaptation of generative flow‑matching manipulation policies using sparse, relative human corrections. During task execution, a human provides brief corrective pose nudges through a lightweight VR interface, and FlowCorrect locally adapts the policy without retraining the backbone, maintaining performance on previously learned scenarios. Experiments on a real‑world robot across four tabletop tasks show that, with a low correction budget, FlowCorrect achieves an 80% success rate on previously failed cases while preserving performance on solved scenarios.
By Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes
arXiv:2609.38059v1 Announce Type: cross
Abstract: Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a sca...
By Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia, Haoyang Huang, Nan Duan, Jiaya Jia
arXiv:2609.16864v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks bec...
By Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain
HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.
By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
arXiv:2608.22296v1 Announce Type: cross
Abstract: Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact th...
By Haoran Lin, Mingyu Yang, Pengfei Qi, Kehan Chen, Qiang Diao, Liangji Zeng, Wenrui Chen, Yaonan Wang, Kailun Yang
Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vision-Language-Action models benefit from large-scale pretraining, their predominantly static pretraining objectives provide limited supervision for physical dynamics and temporal causality, leaving control-relevant knowledge to be learned from downstream robot demonstrations.
arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.
By Beomjun Kim, Seong Hyeon Park, Seunghoon Sim, Seungjun Moon, Sanghyeok Lee, Jinwoo Shin