arXiv Machine Learning
Jun 11

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.

By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy
arXiv AI
Jun 16

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

arXiv:2606. 15631v1 Announce Type: cross Abstract: Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute.

By Jeongeun Park, Juhan Park, Taekyung Kim, Sungjoon Choi, Dongyoon Han, Sangdoo Yun
arXiv AI
Jun 12

SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models

arXiv:2602. 04208v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training.

By Hyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang, Seongwon Cho, Jonghyun Choi
arXiv AI
4d ago

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.

By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang