arXiv Machine Learning By Artur Kuramshin, \"Ozg\"ur Aslan, Cyrus Neary, Glen Berseth

Task Robustness via Re-Labelling Vision-Action Robot Data

Read the original on arXiv Machine Learning →

arXiv:2606. 10918v1 Announce Type: cross Abstract: The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 26

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

The paper introduces Hierarchical Skill Retrieval (HSR), a framework that decomposes a target manipulation task into candidate skill sequences and evaluates each plan for semantic plausibility and skill reliability. HSR combines subtask-level language retrieval with behavior-feature reranking to select demonstrations that are both relevant and compatible with the target task, followed by a two-stage pretraining and finetuning pipeline for policy adaptation. Experiments on the LIBERO benchmark and real-world robot tasks show that HSR improves average success rates by 10.3% and 21.3% over the strongest baseline, demonstrating the effectiveness of structured skill-level retrieval for data-efficient Vision‑Language‑Action adaptation.

By Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
arXiv Machine Learning
Jul 8

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

arXiv:2602. 19710v3 Announce Type: replace-cross Abstract: Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision.

By Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu
arXiv AI
Sep 21

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

KnowDemo is a framework that generates diverse robot demonstrations from human videos by leveraging structured manipulation knowledge. It uses a vision‑language model to extract task requirements and permissible execution variations, then resolves these against target‑scene entities to guide candidate generation and screening before motion planning. The resulting demonstrations feature multimodal behavior, alternative contact strategies, and valid subtask orders, and have been shown to improve planning success and enable sim‑to‑real policy transfer across three tasks.

By Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Sch\"afer, Michael Beetz
arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo