arXiv AI By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu

${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

Read the original on arXiv AI →

The paper introduces ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss in discrete action tokenization for Vision‑Language‑Action models. By decomposing latent action features into multiple heads and assigning independent codebooks to each, the tokenizer expands representational expressivity and improves policy performance. Experiments on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks show superior reconstruction fidelity and higher success rates compared to prior methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
1d ago

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.

By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv Computer Vision
Sep 11

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

IMLE‑VLA replaces the iterative action head in vision‑language‑action policies with a single‑step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). This eliminates multi‑step sampling, boosting inference frequency by 3.67× (55 Hz vs. 15 Hz) and achieving the highest average success rate (98.0 %) on the 40‑task LIBERO benchmark while maintaining robustness under perturbations. Real‑world tests on a Franka Emika Panda show smoother, faster motions and a 3.9×–6.6× reduction in inference time per episode.

By Kian Hosseinkhani (Simon Fraser University), Qinhe Peng (University of Pennsylvania), George Shramko (Simon Fraser University), Mehran Aghabozorgi (Simon Fraser University), Jianing Qian (University of Pennsylvania), Tristan Engst (Simon Fraser University), Alireza Moazeni (Simon Fraser University), Dinesh Jayaraman (University of Pennsylvania), Ke Li (Simon Fraser University, Canada CIFAR AI Chair)