Spotter: Let the Embodied Model Lead, and the VLM Reflect for It
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in...
EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.
DreamAvoid introduces a test‑time dreaming framework for Vision‑Language‑Action models to anticipate and avoid failures during critical manipulation phases. It uses a Dream Trigger to detect critical phases, samples candidate action chunks via an Action Proposer, and evaluates short‑horizon futures with a Dream Evaluator trained on success, failure, and boundary data. Experiments on real‑world and simulated tasks show DreamAvoid improves task success rates, achieving 72.5% success versus 48.8% for the base policy and 54.4% for GPC‑RANK.
VLCP (Vision Language Control Policy) is a training‑free robot manipulation approach that keeps a vision‑language model (VLM) frozen and uses it to generate short Python control functions. Unlike traditional methods that retry a fixed policy, VLCP rewrites the control code every K steps based on multi‑view RGB, proprioceptive state, and state delta, allowing failures to be corrected within the same episode. In a 57‑task MuJoCo/RoboVerse benchmark, VLCP achieves 35.1% pooled success versus 3.5% for a single‑query baseline, with a 27.3% within‑episode recovery rate on failed grasps and efficient token usage.
The paper proposes a method to distill world‑model representations into compact Vision‑Language‑Action (VLA) policies. By adding a single feature‑alignment term during VLA training, a frozen world model’s internal features are cached and the student policy learns to match them, eliminating the need for a generative future‑rolling component. The resulting lightweight policy runs in 32 ms on an RTX 5090, achieving high performance on LIBERO and RoboCasa‑GR1, and transfers effectively to real robotic hardware.
World Action Agent (WAA) is a multi‑agent framework that lets vision‑language models (VLMs) directly pilot robots by operating within a visual action workspace. The workspace provides automatically selected contact views, editable action rehearsals, and in‑view correction to refine decisions before low‑level execution. WAA learns procedural skills from expert videos and human teaching, and its interaction traces can train smaller VLMs, achieving state‑of‑the‑art success on LIBERO‑Pro and improving out‑of‑domain performance on robosuite and Qwen3.5‑9B.