Continuous Conditioning of VLAs with Augmenting EMG and Visual Task Descriptors
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 04510v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the visual and language evidence defining the task.
The paper investigates how visuomotor imitation policies fail when visually similar distractors are introduced, framing the issue as conditional visual grounding where the target changes with manipulation phase and task state. Using Action Chunking with Transformers (ACT), the authors systematically vary color and shape similarity of distractors, pinpointing failures to picking and placement stages. They then test distractor augmentation, phase‑dependent attention regularization, and appearance‑based visual prompting, which together significantly improve robustness in simulation and on a physical UR3e robot, and demonstrate similar improvements on a pretrained vision‑language‑action policy for instrument handling.
arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.
arXiv:2606. 19965v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts.
arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...
arXiv:2608.28967v1 Announce Type: cross Abstract: Electroencephalography (EEG)-based robotic control is commonly formulated as a direct classification problem, in which electrical neural signals are...