WM-VS: Progress-Aligned World Models for Closed-Loop Visual Servoing
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference. whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."
CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.
arXiv:2608.30378v1 Announce Type: cross Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: th...
arXiv:2607. 29235v1 Announce Type: cross Abstract: Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout.
arXiv:2609.37398v1 Announce Type: new Abstract: World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning par...
PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.