arXiv:2606. 19408v1 Announce Type: new Abstract: Latent actions provide a compact interface between action-free video and downstream decision-making, yet existing Latent Action Models (LAMs) force every transition through a fixed-capacity bottleneck.
By Takanori Yoshimoto, Yang Hu, Naruya Kondo, Tatsuya Matsushima
arXiv:2609.39563v1 Announce Type: new
Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes betwee...
By Can Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, Ruirui Li
ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.
By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
ARC‑Bench is a new benchmark that tests whether frozen JEPA‑style latent world models can correctly rank candidate actions by latent distance. The study finds that the assumption of latent rankability fails dramatically in both navigation and manipulation tasks, with the top‑scored actions often being suboptimal. Closed‑loop replanning masks this defect, but reducing replanning frequency reveals the underlying ranking failures.
By Zhengshu Zhang, Zhiyuan Li
arXiv:2609.37297v1 Announce Type: new
Abstract: Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion...
By Zhiyuan Li, Wenyan Yang, Pekka Marttinen, Joni Pajarinen
arXiv:2608. 08982v1 Announce Type: new Abstract: Interactive video world models generate rollouts autoregressively under an action stream, yet they are trained and evaluated almost exclusively on factual prediction.
By Yu Ma, Hongli Shi, Xinran Xu