arXiv Machine Learning By Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon

Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

Read the original on arXiv Machine Learning →

The paper introduces CRAFT, a method for improving compositional generalization in vision‑language‑action (VLA) models. It addresses the issue where models rely on visual shortcuts during fine‑tuning, leading them to execute demonstrated skill combinations that match observations rather than the instructed ones. By training with counterfactual instruction–observation pairs and transferring supervision through reusable skill representations, CRAFT enhances success on unseen skill combinations while preserving performance on demonstrated ones across multiple VLA models and benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
3d ago

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv AI
Aug 26

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

The paper introduces Hierarchical Skill Retrieval (HSR), a framework that decomposes a target manipulation task into candidate skill sequences and evaluates each plan for semantic plausibility and skill reliability. HSR combines subtask-level language retrieval with behavior-feature reranking to select demonstrations that are both relevant and compatible with the target task, followed by a two-stage pretraining and finetuning pipeline for policy adaptation. Experiments on the LIBERO benchmark and real-world robot tasks show that HSR improves average success rates by 10.3% and 21.3% over the strongest baseline, demonstrating the effectiveness of structured skill-level retrieval for data-efficient Vision‑Language‑Action adaptation.

By Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
arXiv Machine Learning
Jun 11

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.

By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy
arXiv AI
Jun 12

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

arXiv:2606. 13578v1 Announce Type: cross Abstract: Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach.

By Baochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen