arXiv Computer Vision

Language-Augmented Video Action Anticipation: Design Fundamentals, Benchmarks, and Open Challenges

arXiv AI
Sep 21

Outcome-Conditioned End-Effector Geometry Across Vision-Language-Action Policies

The study examines how different vision‑language‑action (VLA) policies execute a manipulation task by comparing the geometry of their end‑effectors across 15,000 closed‑loop LIBERO rollouts. By pairing 3,600 configuration‑matched policy executions, the authors find that when both policies succeed, their end‑effector trajectories are much closer (median DTW distance 0.0120 m) than when only one succeeds (0.0380 m), a pattern consistent across all tasks, policy pairs, and nine representations. Even successful executions remain as far from same‑task demonstrations as the demonstrations are from each other, indicating that task‑associated geometry, rather than training data overlap, drives these differences.

By Xingyu Lin, Zhuang Li, Zhongrun Wu, Shouquan Zhou, Dehui Du
arXiv AI
Aug 25

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

The paper introduces Intention Distillation (INDI), a method that injects behavior-level intent into Vision‑Language‑Action (VLA) model decoders by leveraging a frozen teacher vision‑language model to interpret demonstrations. During training, the teacher processes the current observation, instruction, coarse action summary, and execution video, producing a multimodal intent representation that the VLA decoder uses alongside trajectory and execution features to predict actions. Experiments on SimplerEnv‑Bridge, RoboCasa Kitchen, and real‑world tasks show that INDI consistently improves success rates, especially on longer‑horizon tasks, demonstrating that explicit modeling of semantic intent benefits action decoders.

By Sangoh Lee, Sangwoo Mo, Wook-Shin Han
arXiv AI
Sep 1

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

arXiv:2605.30117v2 Announce Type: replace Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VL...

By Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju
arXiv Computer Vision
Sep 24

CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

CereVLA is a cerebellum-inspired framework that enhances vision‑language‑action (VLA) policies by adding lightweight residual refinement and predictive consequence evaluation to frozen VLA execution. It generates corrective actions via flow‑based residual refinement, then assesses their short‑ and interval‑horizon impacts using a recurrent state‑space model and a history‑aware classifier, suppressing unfavorable corrections with a lightweight governor. Experiments on LIBERO‑10, LIBERO‑GOAL, and SO‑101 show that CereVLA improves task success rates and reduces control steps compared to state‑of‑the‑art baselines.

By Shuai Zeng, Yuxuan Liang, Hangmiao Hu, Fobao Zhou, Zixiang Wang, Wenxi Hong, Hang Zhao
arXiv AI
Jun 16

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

arXiv:2606. 15631v1 Announce Type: cross Abstract: Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute.

By Jeongeun Park, Juhan Park, Taekyung Kim, Sungjoon Choi, Dongyoon Han, Sangdoo Yun