CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.
By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
arXiv:2609.39971v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
arXiv:2608.21402v1 Announce Type: cross
Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera...
By Bingqi Huang, Bingchuan Wei, Yingkai Cai, Zhaokui Wang
arXiv:2608.30378v1 Announce Type: cross
Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: th...
By Botong Zhao, Fang Yu, Tim, Senhua Zhu, Xinyuan Chen, Yue Lu
The paper proposes a method to distill world‑model representations into compact Vision‑Language‑Action (VLA) policies. By adding a single feature‑alignment term during VLA training, a frozen world model’s internal features are cached and the student policy learns to match them, eliminating the need for a generative future‑rolling component. The resulting lightweight policy runs in 32 ms on an RTX 5090, achieving high performance on LIBERO and RoboCasa‑GR1, and transfers effectively to real robotic hardware.
By Trung Dao, Sankalp Yamsani, Jaden Park, Joohyung Kim, Yong Jae Lee
ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference.
whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."
By Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2609.39235v1 Announce Type: cross
Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
By Ali Alrasheed, Basim Azam, Naveed Akhtar
arXiv:2606. 01520v1 Announce Type: new Abstract: A single action-conditioned latent predictive architecture can in principle be trained on the structured state of a driving scene, a robot workspace, or a financial order book.
By Shayan Shokri
arXiv:2609.20892v1 Announce Type: cross
Abstract: Closed-loop visual servoing requires predictions that indicate whether an action reduces task error, not only whether the action is plausible. We cal...
By Guanzhong Sun, Junyi Ma, Yixuan Zhou, Yuxuan Wu, Yanzi Miao, Hesheng Wang
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that replaces the Gaussian regularizer used in Joint‑Embedding Predictive Architectures (JEPAs) with a contrastive inverse‑dynamics head. AC‑MTM trains a forward latent‑prediction model while an auxiliary inverse‑dynamics task forces the encoder to distinguish actions from latent transitions, preventing collapse without requiring a target network or reconstruction loss. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM matches or surpasses the performance of the Gaussian‑based SIGReg regularizer, achieving up to 20–24 point improvements on the OGBench Visual Scene benchmark.
By Jack Boylan, Chris Hokamp
arXiv:2606. 03134v1 Announce Type: cross Abstract: Imitation-learning policies for robot manipulation inherit the quality of the success labels attached to their training episodes, and those labels are usually produced by the robot's own success check.
By Aarav Bedi (University of California, Berkeley)