The paper introduces BAS‑VLA, a task‑semantic action calibration framework for vision‑language‑action models that addresses two failure modes: unnecessary action drift under appearance changes and insufficient behavioral change under semantic alterations. BAS‑VLA uses a breaking‑centered calibration core and a selective evidence‑gated preserving auxiliary to maintain performance on clean and semantics‑preserving conditions while suppressing stale‑task behavior. Experiments on OpenPI‑pi0.5 and LIBERO‑Object Milk‑Swap show high success rates on clean and preserved tasks, a dramatic drop under target‑object swaps, and improved robustness to style shifts from 42% to 70% without harming clean performance.
By Shuaijun Liu, Feiyang You, Chengyu Wu, Shuyang Hao, Chenglong Zhang, Jingyao Cai, Xingwei Chen, Ningxin Su
arXiv:2605. 20448v2 Announce Type: replace-cross Abstract: Vision-language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit?
By Animesh Maheshwari, Divyansh Sahu, Nishit Verma
arXiv:2606. 04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes.
By Zirui Wang, Junwei Yu, Adam Yala, David M. Chan, Joseph E. Gonzalez, Trevor Darrell
The paper introduces a progressive training strategy for embodied vision‑language models aimed at reducing spatio‑temporal hallucinations. It first creates a Chain‑of‑Thought dataset that breaks complex reasoning into detailed spatiotemporal steps, then uses supervised pre‑training on this dataset followed by fine‑tuning with weakly‑labeled data. Experiments show the method improves backbone accuracy and narrows the forward‑backward performance gap from over 70% to 6.53%, indicating stronger dynamic reasoning and fewer temporal biases.
By Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue, Menglan Tang, Checheng Yu, Xunzhe Zhou, Sashuai Zhou, Tao Jin, Lixin Yang, Xiangyu Yue, Zhou Zhao
arXiv:2602.06652v2 Announce Type: replace
Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...
By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv:2603. 22278v2 Announce Type: replace-cross Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations.
By Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham