arXiv AI
Jun 30

WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL

arXiv:2602. 13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots.

By Zhennan Jiang, Shangqing Zhou, Yutong Jiang, Zefang Huang, Mingjie Wei, Yuhui Chen, Tianxing Zhou, Zhen Guo, Hao Lin, Quanlu Zhang, Yu Wang, Haoran Li, Chao Yu, Dongbin Zhao
arXiv Computer Vision
Sep 4

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

SV-WAM is a surround‑view world‑action model that keeps all six camera views while enabling efficient inference by discarding the video branch at deployment. It uses future‑video prediction as dense training supervision and an action‑centered causal mask to prevent action tokens from attending to future‑video tokens during joint denoising. A differentiable drivable‑area compliance regularizer penalizes vehicle‑footprint corners near or crossing drivable boundaries, improving safety and boundary awareness. Experiments on NAVSIMv2 and nuScenes show state‑of‑the‑art planning performance with low latency and strong zero‑shot transfer.

By Jinyang Wang, Shiwei Li, Junjian Wang, Zhiqiang Deng, Jianbin Gao, Yihang Zhao, Liu Liu, Yongjia Zhao, Jinlong Chen, Huirui Xu, Yifeng Pan, Kangwei Liu, Fan Ren, Ji Tao, Minghao Yang
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
Hugging Face Trending Papers
Sep 3

SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

SV-WAM is a surround‑view world‑action model that keeps all six camera views for autonomous driving while enabling efficient inference by discarding the video branch during deployment. It uses future‑video prediction as dense training supervision and introduces an action‑centered causal mask to prevent future‑video tokens from influencing action tokens during joint denoising. A differentiable drivable‑area compliance regularizer further improves safety by penalizing vehicle‑footprint corners that approach or cross drivable boundaries. Experiments on NAVSIMv2 and nuScenes show state‑of‑the‑art planning performance with low latency and strong zero‑shot transfer.