arXiv Machine Learning

Auditing Instruction-Trajectory Mismatches in Multimodal Robot Demonstrations

arXiv:2608. 07895v1 Announce Type: cross Abstract: Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction.

arXiv Machine Learning
Jun 5

Is Diversity All You Need for Scalable Robotic Manipulation?

arXiv:2507. 06219v2 Announce Type: replace-cross Abstract: Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood.

By Modi Shi, Li Chen, Jin Chen, Yuxiang Lu, Chiming Liu, Guanghui Ren, Ping Luo, Di Huang, Maoqing Yao, Hongyang Li
arXiv Machine Learning
1d ago

Same Scene, Different Task: Skill Alignment for Compositional Generalization in VLAs

The paper introduces CRAFT, a method for improving compositional generalization in vision‑language‑action (VLA) models. It addresses the issue where models rely on visual shortcuts during fine‑tuning, leading them to execute demonstrated skill combinations that match observations rather than the instructed ones. By training with counterfactual instruction–observation pairs and transferring supervision through reusable skill representations, CRAFT enhances success on unseen skill combinations while preserving performance on demonstrated ones across multiple VLA models and benchmarks.

By Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon
arXiv AI
Aug 26

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

The paper introduces Hierarchical Skill Retrieval (HSR), a framework that decomposes a target manipulation task into candidate skill sequences and evaluates each plan for semantic plausibility and skill reliability. HSR combines subtask-level language retrieval with behavior-feature reranking to select demonstrations that are both relevant and compatible with the target task, followed by a two-stage pretraining and finetuning pipeline for policy adaptation. Experiments on the LIBERO benchmark and real-world robot tasks show that HSR improves average success rates by 10.3% and 21.3% over the strongest baseline, demonstrating the effectiveness of structured skill-level retrieval for data-efficient Vision‑Language‑Action adaptation.

By Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
arXiv Computer Vision
Aug 26

Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning

arXiv:2608.24885v1 Announce Type: cross Abstract: Action-conditioned world models are increasingly used as learned simulators for policy evaluation and improvement, yet their effectiveness rests on a...

By Sixiang Chen, Jiaming Liu, Jixian Wu, Yichen Guo, Tinghao Wang, Siyuan Qian, Hao Chen, Jiajun Cao, Jian Tang, Shanghang Zhang
arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem
arXiv AI
6d ago

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.

By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv AI
Sep 18

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

The paper proposes using action‑similarity supervision to improve cross‑embodiment transfer in latent action models (LAMs). By training the similarity between latent actions to match the similarity of ground‑truth robot action sequences—without predicting the actions themselves—the authors reduce sensitivity to background noise and embodiment differences. Experiments on RoboTwin 2.0 show that this approach more than doubles cross‑embodiment success compared to predicting ground‑truth actions, especially when similarities are computed on end‑effector motion and compared across robots.

By Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo