arXiv Machine Learning

FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning

arXiv:2606. 19408v1 Announce Type: new Abstract: Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed-capacity bottleneck.

arXiv Computer Vision
Sep 25

RotVLA: Rotational Latent Action for Vision-Language-Action Model

RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.

By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu
arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv Computer Vision
Aug 21

What Matters for Latent Actions in Robot Learning

arXiv:2608. 19613v1 Announce Type: cross Abstract: Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions.

By Xizhou Bu, Qingda Hu, Lei Zhou, Lingfeng Zhang, Yingbo Tang, Zihao Liu, Xinyi Tao, Zhiqiang Ma, Qingqiu Huang, Chufeng Tang, Hongbo Wang, Jing Zhang, Jiayi Ma, Hangjun Ye, Wei Li, Xiaoshuai Hao
arXiv Computer Vision
2d ago

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

arXiv:2610.01741v1 Announce Type: new Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, exist...

By Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
arXiv AI
Jul 3

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

arXiv:2607. 01586v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because existing models often differ in architecture, data, action space, and evaluation protocol.

By Guoyang Xia, Fengfa Li, Hongjin Ji, Lei Ren, Fangxiang Feng, Kun Zhan, Yan Xie
arXiv Machine Learning
4d ago

FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales

FlexiWorld is a JEPA-based latent world model that learns variable‑length action chunks across multiple time scales for goal‑directed planning. It jointly trains a causal action encoder and an autoregressive actor, using mixed‑span goal supervision and Student Forcing to reduce exposure bias. In experiments on four benchmarks, FlexiWorld with the Actor‑Residual Cross‑Entropy Method (ARCEM) achieves higher mean success rates than the strongest baseline and supports flexible planning chunk lengths without retraining.

By Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun, Jiaqi Wang, Damien Scieur, Yunze Liu