arXiv:2609.15005v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) policies perform robot manipulation tasks using multimodal inputs such as visual observations, proprioceptive states, an...
By Jinwoong Kim, Sangjin Park
The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.
By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
arXiv:2608. 04510v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the visual and language evidence defining the task.
By Suhas Hegde, Jitendra Yasaswi Bharadwaj Katta
arXiv:2606. 15631v1 Announce Type: cross Abstract: Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute.
By Jeongeun Park, Juhan Park, Taekyung Kim, Sungjoon Choi, Dongyoon Han, Sangdoo Yun
The paper argues that diffusion-based action policies can use a frozen, observation‑free backbone as a reusable trajectory prior, with task adaptation handled entirely by the conditioning pathway. By pretraining a general action head on forward‑kinematics data and then freezing it, the authors show that a single backbone can match or outperform normally trained models on MimicGen and LIBERO. Their experiments reveal that a small 5 M‑parameter MLP backbone can rival large U‑Net and transformer backbones, indicating that action backbones are often over‑parameterized and that image‑style architectures may not be the best fit for low‑dimensional action generation.
By Jian Zhou, Sihao Lin, Shuai Fu, Zerui Li, Gengze Zhou, Qi WU
arXiv:2609.36118v1 Announce Type: new
Abstract: Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone la...
By Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames, Yisong Yue
ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.
By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv:2607. 16506v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometrically successful for the current skill can be brittle for downstream skills.
By Yuhan Liu, Xinyu Zhang, Litao Liu, Abdeslam Boularias
arXiv:2609.37165v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific f...
By Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, Gi-Cheon Kang, Byoung-Tak Zhang
arXiv:2608. 04692v1 Announce Type: cross Abstract: Task-vector arithmetic offers a closed-form way to modify a model, yet its behavioral locality remains unclear in closed-loop robot control.
By Shaoguang Wang, Weiyu Guo, Rushi Dai, Yiren Zhao, Yandong Guo, Hui Xiong
arXiv:2608.30378v1 Announce Type: cross
Abstract: Direct vision-language-action policies generate continuous robot actions efficiently, but standard behavior cloning leaves two complementary gaps: th...
By Botong Zhao, Fang Yu, Tim, Senhua Zhu, Xinyuan Chen, Yue Lu
arXiv:2609.12081v1 Announce Type: cross
Abstract: Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain...
By Guangyu Chen, Qiwei Liang, Shaolong Zhu, Tianxing Chen, Zikuan Xiao, Yifan Xie, Lingfeng Zhang, Ping Luo, Renjing Xu, Wenbo Ding