arXiv AI

MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding

arXiv:2606. 00174v1 Announce Type: cross Abstract: Electromyography (EMG) directly reflects muscle activation and is a key sensing modality for gesture recognition, prosthetic control, and wearable interaction.

arXiv Machine Learning
Jun 24

Lightweight Test-Time Adaptation for EMG-Based Gesture Recognition

arXiv:2601. 04181v2 Announce Type: replace Abstract: Reliable long-term decoding of gestures from surface electromyography (EMG) is hindered by signal drift caused by electrode displacement, muscle fatigue, and/or posture changes.

By Nia Touko, Matthew O A Ellis, Cristiano Capone, Alessio Burrello, Elisa Donati, Luca Manneschi
arXiv Computer Vision
Sep 21

SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

SignGPT is a unified, pose‑based framework that performs gloss‑free sign language translation (SLT) and generation (SLG) by integrating part‑aware hierarchical representations of body, hand, and facial motion into a shared language model. It uses asymmetric multi‑token prediction and progressive training for bidirectional modeling, and is evaluated on How2Sign (ASL) and Phoenix‑2014T (DGS) with benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers demonstrates a sign‑to‑sign response pipeline, suggesting that unified modeling can support sign language conversation (SLC).

By Ronghui Li, Jun Dong, Zhongyuan Hu, Zunnan Xu, Jun Zhou, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Jie Guo, Xiu Li, Linchao Bao
arXiv Computer Vision
Sep 3

DuoGesture: Motion-Grounded Semantic Conditioning and Biomechanical Beat Priors for Co-Speech Gesture Generation

DuoGesture is a co‑speech gesture generation model that separates gesture synthesis into a semantic stream and a beat stream, coordinated by a Semantic Variational Information Bottleneck that decides when semantic gestures override rhythmic motion. The semantic stream uses Motion‑Grounded Semantic Conditioning, replacing word embeddings with motion‑language representations to provide motion‑aligned semantic priors for rare gesture triggers. The beat stream is regularised by an Inertial Beat Prior, an anthropometry‑weighted arm‑chain module that reduces jitter and improves rhythmic consistency. Experiments show DuoGesture outperforms strong baselines and ablations confirm the complementary roles of semantic grounding, stochastic stream selection, and biomechanical regularisation.

By Ferdinand Paar, Lanmiao Liu, Asl{\i} \"Ozy\"urek, Serge Thill, Esam Ghaleb
arXiv Machine Learning
Aug 27

MyoMechanix: Biomechanically-Grounded Compositional Skilled Activity Understanding and Coaching

MyoMechanix is a multimodal dataset and framework for action quality assessment that incorporates muscle activity and other physiological signals alongside visual data. It contains over 7,500 samples of 20 weight‑loaded actions from 38 subjects, with synchronized RGB video, 3D pose, sEMG, and additional signals. The accompanying Fitness Knowledge Graph structures expert annotations into relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment through the CUBIST engine. The project also introduces MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task, demonstrating that multimodal sensing and structured representations improve performance, interpretability, and error attribution.

By Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
arXiv Computer Vision
Sep 17

Pose2Muscle: Structured Spatio-Temporal Decoding for Discrete Muscle Activity Estimation from Human Pose

Pose2Muscle is a pose-driven framework that estimates discrete muscle activity states without requiring surface electromyography (sEMG) during inference. It reformulates muscle estimation as a structured prediction problem, using multi-scale spatio-temporal attention and a directed acyclic graph-based decoder to capture motion patterns and maintain multiple candidate hypotheses. The authors introduce the PoseEMG-43 dataset, comprising 2,992 movement instances from 43 daily-life actions performed by 14 participants, and demonstrate that Pose2Muscle outperforms baseline methods with high accuracy and correlation metrics.

By Yuepeng Chen, Jiehong Shi, Kaili Zheng, Boyi Zhang, Chenyi Guo, Ji Wu, Xiangling Fu