arXiv AI By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

Read the original on arXiv AI →

The paper introduces M2Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss for continuous action signals by decomposing latent features into multiple heads and assigning independent codebooks to each. This design expands representational expressivity, leading to lower reconstruction error and higher success rates in Vision‑Language‑Action models evaluated on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks. Ablation studies confirm the effectiveness of both multi‑head and multi‑codebook mechanisms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

The paper introduces ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss in discrete action tokenization for Vision‑Language‑Action models. By decomposing latent action features into multiple heads and assigning independent codebooks to each, the tokenizer expands representational expressivity and improves policy performance. Experiments on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks show superior reconstruction fidelity and higher success rates compared to prior methods.

By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu
arXiv AI
Sep 17

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.

By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv Computer Vision
6d ago

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.

By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic