Hugging Face Trending Papers

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by the high cost of action annotations and the heterogeneity of action spaces across platforms.

arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
arXiv Computer Vision
Sep 16

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

GeoLAM is a framework that learns geometry‑grounded latent actions from unlabeled human videos. It uses future‑frame reconstruction with a frozen geometric feature hierarchy and motion supervision from a 4D geometry teacher to capture 3D displacement, image‑plane motion, and surface‑orientation changes. After pretraining, the representation serves as transition targets for a world‑action model trained on robot demonstrations, enabling denoised latent actions and executable action chunks without requiring hand‑pose annotations or future‑video generation during deployment.

By Yifan Xie, Hekun Tian, Jinkun Liu, YuAn Wang, Qiao Sun, Wenbo Ding
arXiv Computer Vision
Sep 3

Spatially Aware World Action Model via Geometric Latent Diffusion

The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.

By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv Machine Learning
Jun 15

$\mu_0$: A Scalable 3D Interaction-Trace World Model

arXiv:2606. 13769v1 Announce Type: cross Abstract: World models that capture how actions induce physical change enable scalable robot learning without reliance on embodiment-specific action labels.

By Seungjae Lee, Yoonkyo Jung, Jusuk Lee, Jonghun Shin, Amir Hossein Shahidzadeh, Yao-Chih Lee, H. Jin Kim, Jia-Bin Huang, Furong Huang
arXiv AI
Jun 16

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.

By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu