arXiv Machine Learning By Ahmed Karim, Leon Chlon

Robot World Models Are Not Invariant to How the Actions Are Written

Read the original on arXiv Machine Learning →

A robot policy trained with either absolute joint targets or delta‑relative actions inherits the chosen action parameterization in its world model, leading to catastrophic failures when the model is exposed to the alternate encoding. Experiments on three robot datasets and two morphologies show retrieval performance drops 2.6–13.4×, goal‑conditioned action selection plummets from 53% to 15%, and predictions for the same future become nearly orthogonal. The issue is not a loss of information—both encodings are highly reconstructible—but a lack of invariance in the action channel, which existing visual‑model invariance research does not address. Averaging over the two encodings partially restores performance, yet the worst‑case disagreement remains high, indicating that the defect persists in certain scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
4d ago

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.

By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
arXiv Computer Vision
Sep 22

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

The paper proposes a method to distill world‑model representations into compact Vision‑Language‑Action (VLA) policies. By adding a single feature‑alignment term during VLA training, a frozen world model’s internal features are cached and the student policy learns to match them, eliminating the need for a generative future‑rolling component. The resulting lightweight policy runs in 32 ms on an RTX 5090, achieving high performance on LIBERO and RoboCasa‑GR1, and transfers effectively to real robotic hardware.

By Trung Dao, Sankalp Yamsani, Jaden Park, Joohyung Kim, Yong Jae Lee
arXiv AI
Aug 24

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference. whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."

By Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang