arXiv AI

Semantic Anchoring for Robotic Action Representations

arXiv:2607. 13597v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot demonstrations degrades this structure and undermines generalization.

arXiv AI
Aug 25

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

The paper introduces Intention Distillation (INDI), a method that injects behavior-level intent into Vision‑Language‑Action (VLA) model decoders by leveraging a frozen teacher vision‑language model to interpret demonstrations. During training, the teacher processes the current observation, instruction, coarse action summary, and execution video, producing a multimodal intent representation that the VLA decoder uses alongside trajectory and execution features to predict actions. Experiments on SimplerEnv‑Bridge, RoboCasa Kitchen, and real‑world tasks show that INDI consistently improves success rates, especially on longer‑horizon tasks, demonstrating that explicit modeling of semantic intent benefits action decoders.

By Sangoh Lee, Sangwoo Mo, Wook-Shin Han
arXiv Computer Vision
3d ago

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv Computer Vision
Aug 27

V-Link: Recovering Lost Visual Representations in Action DiT for Vision-Language-Action Models

V-Link is a method designed to enhance Vision‑Language‑Action (VLA) models by recovering visual representations during the transfer from vision‑language (VL) features to action (A) features. It introduces complementary Spatial and Semantic Query representations that are injected into Action DiT through asymmetric pathways, providing both semantic augmentation and dedicated geometric conditioning for action generation. Experiments on LIBERO, LIBERO‑Plus, RoboTwin 2.0, and real‑world AGIBOT A3 Ultra tasks show significant performance gains over the base GR00T N1.6 model.

By Yehao Lu, Jiarui Yang, Yuning Su, Yufeng Xie, Yu Zhong, Yazhou Zhang, Haiyu Lan, Kaixiang Lu, Peiwen Lin, Chuang Wang, Zequn Qin, Enyu Li, Xi Li
arXiv AI
Sep 17

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

The paper introduces CSWAM, a Causal Semantic World Action Model that enhances FastWAM by integrating a causal semantic expert based on V-JEPA 2.1. This expert provides temporally grounded, appearance‑agnostic representations of semantic state changes and motion, leveraging sparse observation history and causal attention to improve action‑only inference. Experiments on simulation and real‑robot tasks show that CSWAM significantly boosts out‑of‑distribution generalization, raising success rates from 10.16% to 45.18% on RoboTwin 2.0 and from 27.5% to 70.0% across real‑robot tasks.

By Tianbin Liu, Jian Zhu, Taiyi Su, Jianjun Zhang, Chong Ma, Zitai Huang, Yi Xu
arXiv AI
Jul 28

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.

By Yi Liu, Sukai Wang, Dafeng Wei, Xiaowei Cai, Linqing Zhong, Jiange Yang, Guanghui Ren, Jinyu Zhang, Maoqing Yao, Chuankang Li, Xindong He, Liliang Chen, Jianlan Luo