The paper introduces ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss in discrete action tokenization for Vision‑Language‑Action models. By decomposing latent action features into multiple heads and assigning independent codebooks to each, the tokenizer expands representational expressivity and improves policy performance. Experiments on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks show superior reconstruction fidelity and higher success rates compared to prior methods.
By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu
arXiv:2606. 30113v1 Announce Type: cross Abstract: Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actions from discrete codes remains challenging.
By Tengyue Jiang, Chunpu Xu, Jiayue Kang, Yao Mu
arXiv:2606. 14752v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control.
By Xirui Kang, Yanpei Shi, Lucy Liang, Roy Gan, Dongxiu Liu, Pushi Zhang, Danpeng Chen, Xiaoyi Qin, Yinan Zheng, Jinliang Zheng, Hao Wang, Xianyuan Zhan, Hang Su
arXiv:2608. 10484v1 Announce Type: cross Abstract: Action verbs describe not only the physical outcomes of actions, but also how those actions are performed.
By Li Wenjie, Yash Jangir, Ignacy Stepka, Yash Agarwal, Marion Kipsang, Yonatan Bisk
ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.
By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.
By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic
The paper introduces LaPla, a Vision‑Language‑Action framework that uses a latent‑aligned planning approach to convert discrete semantic reasoning into continuous, physics‑constrained driving actions. It employs a residual VQ‑VAE to encode vehicle kinematics into a structured latent space, then projects multimodal inputs—images, past actions, and text—directly into this latent space, allowing a frozen decoder to generate physically plausible trajectories without quantization errors. Experiments on nuScenes and NVIDIA AlpaSim show LaPla reduces long‑horizon L2 error by 15.52% and improves closed‑loop success rates by 33.34 percentage points while cutting inference latency.
By Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
The paper investigates how different action tokenization methods affect closed‑loop control in autoregressive vision‑language‑action models. It compares analytical, linear, and nonlinear representations, showing that lower reconstruction error does not guarantee better policy performance. The study highlights the need to evaluate tokenization on multiple criteria, including sequence predictability and decoder stability, rather than relying solely on reconstruction fidelity.
By Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu
arXiv:2607. 06370v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations.
By Ryuji Oi, Hikari Otsuka, Kosuke Matsushima, Yuki Ichikawa, Masato Motomura, Tatsuya Kaneko, Daichi Fujiki
arXiv:2607. 21670v1 Announce Type: cross Abstract: Action tokenization maps continuous robot action chunks to discrete tokens and has become an important interface for modern visuomotor policies.
By Chaoqi Liu, Yue Zhao, Haonan Chen, Xiaoshen Han, Jiawei Gao, Ehsan Adeli, Yilun Du
The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.
By Ahad Jawaid, Yu Xiang
Vision-language-action (VLA) models commonly adopt an LLM-centric $V \to L \to A$ pathway, where visual observations are projected into the representation space of a large language model before being decoded into robot actions. Although effective, this design incurs substantial computation and memory overhead at every policy invocation.