Hugging Face Trending Papers

Cross-Embodiment Transfer via Behavior-Aligned Representations

Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging.

arXiv AI
Sep 21

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

AtomEgo investigates how to integrate large-scale egocentric human interaction data into embodied foundation model pre‑training. The study uses a curated 2,659‑hour corpus and a scalable data pipeline to evaluate three co‑training paradigms across vision‑language‑action and world‑action architectures. Results show that the benefit of egocentric data depends on both its scale and the quality of alignment with robotic embodiment, offering practical guidance for scalable ego‑robot pre‑training.

By Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie, Xiaoquan Sun, Junyang Zheng, Zhuoyang Song, Jiaxing Zhang, Jiayu Chen
arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
arXiv AI
Aug 26

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

The paper introduces Hierarchical Skill Retrieval (HSR), a framework that decomposes a target manipulation task into candidate skill sequences and evaluates each plan for semantic plausibility and skill reliability. HSR combines subtask-level language retrieval with behavior-feature reranking to select demonstrations that are both relevant and compatible with the target task, followed by a two-stage pretraining and finetuning pipeline for policy adaptation. Experiments on the LIBERO benchmark and real-world robot tasks show that HSR improves average success rates by 10.3% and 21.3% over the strongest baseline, demonstrating the effectiveness of structured skill-level retrieval for data-efficient Vision‑Language‑Action adaptation.

By Haoran Hao, Shahram Najam Syed, Jeff Schneider, Jeffrey Ichnowski
arXiv Machine Learning
Aug 20

The Embodiment Gap in Robot Foundation Models

The paper discusses the "embodiment gap" in robot foundation models, highlighting that while models can generalize across tasks, additional work is often needed to deploy them on specific robot bodies. It surveys what components can be reused across different robot embodiments and what must be implemented anew, mapping existing methods along axes of shared structure and adaptation stage. The authors propose a reporting framework to better assess adaptation efforts and identify remaining challenges for cross-embodiment learning.

By Yukiyasu Domae, Keisuke Shirai, Hanbit Oh, Ryoichi Nakajo, Tomohiro Motoda, Koshi Makihara, Masaki Murooka, Takuma Yagi, Yoshiaki Bando, Ryo Hanai
arXiv AI
Sep 18

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

The paper proposes using action‑similarity supervision to improve cross‑embodiment transfer in latent action models (LAMs). By training the similarity between latent actions to match the similarity of ground‑truth robot action sequences—without predicting the actions themselves—the authors reduce sensitivity to background noise and embodiment differences. Experiments on RoboTwin 2.0 show that this approach more than doubles cross‑embodiment success compared to predicting ground‑truth actions, especially when similarities are computed on end‑effector motion and compared across robots.

By Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo
arXiv Computer Vision
Aug 31

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

The paper introduces VLAct, a Vision‑Language‑Action model that focuses on representation‑centric continued pre‑training rather than merely scaling robot data. VLAct is trained on diverse, multi‑embodiment robot data and preserves a broad VLM prior while encouraging shared action semantics across embodiments. Experiments across simulation, real‑world, and unseen‑embodiment settings show that VLAct consistently outperforms existing industrial VLA systems, achieving high success rates with only a modest compute budget and open‑source data.

By Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia