arXiv AI By Liuhaichen Yang, Zhengyang Zhong, Hanshang Zhu, Ningwei Bai, Qichen Yin, Zhi Han, Jiarui Qin, Zhuang Jiang, Chenchao Sheng, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Bo Liu, Yi Dong, Zezhi Tang

WAM-OPD: Joint Video-Action Supervision for World Action Model Post-Training with On-Policy Distillation

Read the original on arXiv AI →

WAM-OPD introduces a method for post‑training improvement of World Action Models (WAMs) by collecting student rollout histories and querying a stronger teacher for paired video and action targets. The student learns from both modalities while preserving its one‑step generation capability at deployment. Experiments on 12 RoboTwin 2.0 tasks and four real‑robot tasks show success rates rising from 33.8% to 65.7% and from 51.4% to 64.6%, respectively, with joint video‑action supervision outperforming single‑modality approaches.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
5d ago

EVO-WAM: Evolving World Action Models through Video-Action Verification

arXiv:2609.38057v1 Announce Type: new Abstract: Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action m...

By Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, Zhuotao Tian
arXiv Computer Vision
23h ago

Native Action-Prior Learning from Videos for World Action Models

arXiv:2610.03391v1 Announce Type: new Abstract: World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotat...

By Zhaochong An, Fei Zhang, Menglin Jia, Duncan Frost, Zijian Zhou, Yikai Wang, Xudong Wang, Aditya Patel, Belinda Zeng, Tao Xiang, Serge Belongie, Amir Bar, Sen He
arXiv Computer Vision
Sep 21

MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

MT‑WAM enhances the Fast‑WAM framework by adding complementary supervision for future 2‑D point trajectories and visual features while keeping the original training objectives. A lightweight dual‑stream branch and structured attention mask isolate motion‑specific processing, and motion‑stream tokens provide additional dynamics cues to the action expert. During inference, MT‑WAM skips future‑video prediction, using cached video and motion information to achieve higher success rates on LIBERO, LIBERO‑Plus, RoboTwin 2.0 Clean2Rand, and several real‑world tasks.

By Yiguang Yang, Jiankun Peng, Xiaoming Wang, Yiran Zhang, Zhibo Fang
arXiv Computer Vision
Aug 27

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.

By Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu