arXiv Machine Learning

Robotic Policy Adaptation via Weight-Space Meta-Learning

arXiv:2606. 07217v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from large corpora of demonstrations and action labels.

arXiv AI
Sep 18

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

HIL-UMI is a policy-guided Universal Manipulation Interface that enables robot‑free, human‑in‑the‑loop post‑training of vision‑language‑action models. By querying the current policy during handheld demonstrations and using an Energy Score to detect out‑of‑distribution states, it selectively collects new data and refines a progress‑based advantage estimator. The updated estimator then drives advantage‑conditioned behavioral cloning, improving performance on long‑horizon and precise manipulation tasks while reducing per‑frame collection time compared to HG‑DAgger.

By Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong
arXiv AI
2d ago

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

arXiv:2609.36416v2 Announce Type: replace-cross Abstract: Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isol...

By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Thomas Wolf, Jackson Lee, Pragna Mannam
arXiv AI
4d ago

FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

arXiv:2609.36416v1 Announce Type: cross Abstract: Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions....

By Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, Jackson Lee, Thomas Wolf, Pragna Mannam
arXiv Machine Learning
Sep 21

From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

The paper introduces PARTS, a real‑world subtask reinforcement learning framework that fine‑tunes a pretrained robot policy by focusing on critical bottleneck subtasks while keeping the base policy frozen. It uses agent‑generated selectors and success verifiers to provide local rewards, enabling learning even when full‑task successes are rare. Experiments on bimanual YAM and single‑arm Franka robots show that PARTS raises complete‑task success from 32% to 61% and from 50% to 95%, respectively, with only tens of minutes of real‑world RL rollouts and minimal human intervention.

By Sichang Su, Benjamin Yang, Zhiyun Deng, Boyuan Liang, Yip Fun Yeung, Zelin Wang, Lingfeng Sun
arXiv AI
Aug 26

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

The paper introduces Hierarchical Skill Retrieval (HSR), a framework that decomposes a target manipulation task into candidate skill sequences and evaluates each plan for semantic plausibility and skill reliability. HSR combines subtask-level language retrieval with behavior-feature reranking to select demonstrations that are both relevant and compatible with the target task, followed by a two-stage pretraining and finetuning pipeline for policy adaptation. Experiments on the LIBERO benchmark and real-world robot tasks show that HSR improves average success rates by 10.3% and 21.3% over the strongest baseline, demonstrating the effectiveness of structured skill-level retrieval for data-efficient Vision‑Language‑Action adaptation.

By Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
arXiv Computer Vision
3d ago

Correcting WHERE, Preserving HOW: Compositional Generalization for Vision-Language-Action Models via Referential Guidance

arXiv:2609.38616v1 Announce Type: cross Abstract: While Vision-Language-Action (VLA) models enable flexible action generation, their generalization across diverse environmental elements, including ma...

By Yanyan Zhang, Disheng Liu, Xinpeng Li, Chaoda Song, Mohsen Hariri, Debargha Ganguly, Wang Yang, Kai Ye, Bryce Grant, Vipin Chaudhary, Yu Yin
arXiv AI
Aug 19

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

EXPO-FT is a system that enables stable, sample‑efficient reinforcement learning fine‑tuning of pretrained Vision‑Language‑Action (VLA) policies. It achieves perfect success on a range of manipulation tasks—such as routing string lights, striking a pool ball, and inserting a flower into a wine bottle—using only about 19.1 minutes of online robot data. The approach outperforms both RL-from-scratch and existing VLA fine‑tuning methods, and the authors provide an open‑source codebase to support wider adoption.

By Perry Dong, Kuo-Han Hung, Tian Gao, Dorsa Sadigh, Chelsea Finn
arXiv AI
Jun 16

Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time

arXiv:2606. 15631v1 Announce Type: cross Abstract: Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fine-tuning, making adaptation costly in both data collection and compute.

By Jeongeun Park, Juhan Park, Taekyung Kim, Sungjoon Choi, Dongyoon Han, Sangdoo Yun