arXiv AI By Haoyu Gu, Haotian Lu, Jingrun Du, Xiao-Ping Zhang

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units

Read the original on arXiv AI →

arXiv:2608. 03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 4

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.

By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
arXiv AI
Sep 17

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.

By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv Computer Vision
Sep 25

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.

By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic