Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation
arXiv:2608. 19490v1 Announce Type: cross Abstract: State-of-the-art vision-language-action (VLA) models such as $\pi_{0.
arXiv:2608. 19490v1 Announce Type: cross Abstract: State-of-the-art vision-language-action (VLA) models such as $\pi_{0.
Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support t...
arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.
arXiv:2606. 25800v1 Announce Type: new Abstract: Effective online adaptation of vision-language-action (VLA) models remains challenging, as sparse rewards provide weak supervision for high-dimensional autoregressive action policies.
arXiv:2512. 20014v3 Announce Type: replace-cross Abstract: While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring my cup," where the robot must act on one specific instance among visually similar objects.
arXiv:2609.36416v1 Announce Type: cross Abstract: Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions....
Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.
arXiv:2607. 02466v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions, and actions that are costly to collect at scale.
arXiv:2609.36416v2 Announce Type: replace-cross Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isol...
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task ca...
arXiv:2606. 07217v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels.
arXiv:2608. 11363v1 Announce Type: cross Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction.