arXiv AI

HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models

arXiv:2607. 04265v1 Announce Type: cross Abstract: World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion.

arXiv AI
Sep 18

JEPA-WAM: Connecting Generated Visual Instructions to World Action Models through JEPA Latent Representations

JEPA-WAM enhances World Action Models (WAMs) by pairing text instructions with stochastically generated visual cues, using a text-to-image generator and a frozen V‑JEPA encoder to create dense goal representations. These representations are compressed into goal tokens that condition both video and action experts via cross‑attention, enabling the model to better ground instructions. On a new real‑robot benchmark, JEPA‑WAM attains 87.3%, 74.5%, and 80.9% success rates across in‑distribution, out‑of‑distribution scenes, and out‑of‑distribution instructions, outperforming prior methods by significant margins.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv Machine Learning
Sep 21

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

The paper introduces PARTS, a real‑world subtask reinforcement learning framework that fine‑tunes a pretrained robot policy by focusing on critical bottleneck subtasks while keeping the base policy frozen. It uses agent‑generated selectors and success verifiers to provide local rewards, enabling learning even when full‑task successes are rare. Experiments on bimanual YAM and single‑arm Franka robots show that PARTS raises complete‑task success from 32% to 61% and from 50% to 95%, respectively, with only tens of minutes of real‑world RL rollouts and minimal human intervention.

By Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun
arXiv Computer Vision
4d ago

EVO-WAM: Evolving World Action Models through Video-Action Verification

arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...

By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
arXiv AI
Sep 24

BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

arXiv:2609.27450v1 Announce Type: cross Abstract: Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale er...

By Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao
arXiv AI
Sep 18

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.

By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
arXiv Machine Learning
Sep 25

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

The paper introduces Predictive Action Chunk Learning (PACL), a method for improving robot manipulation policies using mixed-quality deployment experience. PACL first trains a predictive chunk-level critic to evaluate temporally extended action sequences, then uses the critic’s quality estimates to guide a diffusion actor that learns from both successful and failed rollouts. Experiments on simulated and real robots demonstrate that PACL consistently enhances pretrained policies and outperforms strong imitation learning and offline reinforcement learning baselines.

By Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, Chen Lv
arXiv AI
Jun 9

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

arXiv:2606. 09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning.

By Jisong Cai, Long Ling, Shiwei Chu, Zhongshan Liu, Jiayue Kang, Zhixuan Liang, Wenjie Xu, Yinan Mao, Weinan Zhang, Xiaokang Yang, Ru Ying, Ran Zheng, Yao Mu