Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow ne...
arXiv:2608.23526v1 Announce Type: new
Abstract: World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pendulum video...
By Richard Bao
arXiv:2609.32551v2 Announce Type: replace-cross
Abstract: Text-conditioned human-object interaction (HOI) generation requires body motion, object trajectories & rotations, and hand articulation to re...
By Dawei Guan, Di Yang, Jiangtao Wang
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2607. 29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout.
By Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang
The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.
By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
The paper introduces Action Forcing, a method that transforms ordinary unlabeled video into action‑supervised training data by extracting egomotion bases through principal component analysis of pixel displacements. This approach yields grounded throttle–yaw control signals without requiring instrumented platforms or manual annotation, and it trains a high‑capacity video model while preventing pixel‑level overfitting via an online latent critic. The authors also critique standard video generation metrics and propose a reference‑free evaluation that measures controllability, plausibility, conjuring, and geometric integrity, showing that their model can reverse, scale, and compose actions despite limited reverse‑action data.
By Ashish Sundar, Tiankuo Hou, Zhong Fan, Chunbo Luo, Xiaoyang Wang
arXiv:2608.24855v1 Announce Type: new
Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
By Hsiang-Wei Huang, Jianxu Shangguan, Junbin Lu, Jenq-Neng Hwang
RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.
By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
By Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
arXiv:2608.10860v3 Announce Type: replace-cross
Abstract: World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel recons...
By Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox