arXiv AI
Sep 7

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

The paper investigates how visuomotor imitation policies fail when visually similar distractors are introduced, framing the issue as conditional visual grounding where the target changes with manipulation phase and task state. Using Action Chunking with Transformers (ACT), the authors systematically vary color and shape similarity of distractors, pinpointing failures to picking and placement stages. They then test distractor augmentation, phase‑dependent attention regularization, and appearance‑based visual prompting, which together significantly improve robustness in simulation and on a physical UR3e robot, and demonstrate similar improvements on a pretrained vision‑language‑action policy for instrument handling.

By Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger
arXiv AI
Jun 3

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.

By Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
arXiv Computer Vision
Oct 1

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin