arXiv:2606. 10683v1 Announce Type: cross Abstract: Dexterous hands are essential for fine-grained manipulation, but their hardware designs vary substantially across embodiments.
By Dong Fang, Youjun Wu, Yuanxin Zhong, Rui Zhang, Yunlong Wang, Xiaosong Jia, Yu-Gang Jiang
M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.
By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
arXiv:2609.13269v1 Announce Type: cross
Abstract: Gesture recognition on video is normally posed as classification: label each frame, then act on the label. That is adequate for control, where a comm...
By Amey Thakur
ActionPiece rethinks how actions are tokenized for autoregressive vision‑language‑action models by introducing physical rank consistency (PRC) to evaluate relational fidelity of reconstructed actions. The method jointly supervises representation learning and quantization to preserve local physical distance rankings, improving both PRC and policy success. Experiments on LIBERO, LIBERO‑Plus, SimplerEnv, and VLA‑Arena show significant gains over baseline tokenizers.
By Shijie Lian, Bin Yu, Zhaolong Shen, Xiaopeng Lin, Yichao Du, Zhirui Zhang, Laurence T. Yang, Kai Chen
arXiv:2606. 18092v1 Announce Type: cross Abstract: Cross-end-effector grasp generation seeks a unified model that generalizes across objects and across embodiments ranging from parallel grippers to dexterous end effectors.
By Wanhao Niu, Qiyan Ke, Yuan Sun, Hao Sun, Jie Xu, Muyuan Ma, Ruiqi Hu, Fuchun Sun
The paper introduces Direction-Scale Decomposition (DSD), an action representation that separates translation and rotation increments into direction and scale components before tokenization. DSD is evaluated with uniform binning and a B-spline tokenizer (BEAST) in both simulation and real-world manipulation tasks, showing improved success rates on LIBERO and SimplerEnv, especially under mixed-dataset training. Real-robot experiments confirm performance gains with and without robotics pretraining, supporting DSD as an effective representation for discrete-token vision-language-action models.
By Yufei Duan, Hang Yin, Alberta Longhini, Chao Tang, Danica Kragic