arXiv Machine Learning By Weisen Zhao, Lam Nguyen, Zhicong Lu, Yuzhang Shang

C$^3$ache: Accelerating World Action Models with Cross Inference Chunk Cache

Read the original on arXiv Machine Learning →

arXiv:2606. 08962v1 Announce Type: new Abstract: World Action Models (WAMs) generalize better than standard Vision-Language-Action (VLA) policies to novel motions and environments, because a video-modeling objective lets them learn from abundant unlabeled video rather than scarce labeled robot demonstrations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 25

Rolling-WAM: World Action Models with Rolling Imagination

Rolling-WAM is a new formulation for World Action Models that spreads the joint video-action denoising process across multiple replanning cycles. It keeps a sliding window of video-action chunks at different noise levels, fully denoising the immediate chunk for execution while partially refining future chunks. This approach reduces latency, improves closed-loop responsiveness, and achieves a 4.5× speedup in steady-state replanning compared to standard WAMs while maintaining competitive manipulation performance.

By Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang