arXiv Computer Vision

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.

arXiv AI
Aug 11

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

arXiv:2608. 09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress.

By Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie
arXiv Machine Learning
2d ago

EMODY Flow: Emotion-Aware Audio-Driven Full-Body Motion Generation

EMODY Flow is a lightweight flow‑matching framework that generates synchronized full‑body motion and facial expressions conditioned on speech and emotion. It attaches to a frozen Qwen‑3 Omni model, reusing its audio codecs to drive two parallel DiT generators for SMPL‑X body pose and FLAME facial expressions. An auxiliary emotion classifier at training time restores emotion sensitivity, enabling EMODY Flow to achieve state‑of‑the‑art gesture quality on BEAT2 and zero‑shot facial animation on TFHP, with significant improvements in FGD, Beat Correlation, and Diversity metrics.

By Harsh Kumar Agarwal, Xavier Alameda-Pineda, Olivier Perrotin
arXiv Computation and Language
Sep 3

SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation

SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.

By Marshall Thomas, Edward Fish, Richard Bowden
arXiv Computer Vision
3d ago

SignMimic: Robust High-Quality Sign Language Motion Generation via Human-Shape-Oblivious Pose Transfer Guidance

arXiv:2609.14122v1 Announce Type: new Abstract: We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer...

By Zhewen He (New York University Abu Dhabi), Junyi Yu (New York University Abu Dhabi), Haomian Huang (New York University Abu Dhabi), Zhenhua Li (ChatSign Technology), Yi Fang (New York University Abu Dhabi, ChatSign Technology)
arXiv Computer Vision
Aug 31

SignRR: Retrieve and Refine Real Motion for Sign Language Production

SignRR is a new sign language production framework that combines retrieval of real sign motion segments with a learned refinement step to produce globally coherent signing sequences. It starts from a dictionary of authentic sign segments and refines them using a part-aware Residual VQ‑VAE, preserving fine hand articulation while handling temporal length differences in latent space. Experiments on PHOENIX14T and CSL‑Daily demonstrate state‑of‑the‑art back‑translation performance and competitive pose quality.

By Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan, Gissella Bejarano
arXiv Machine Learning
Jun 11

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.

By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv Computer Vision
Sep 3

SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

SignMatch introduces a prototype‑structured embedding space that learns to match dictionary sign videos with continuous sign language footage based solely on visual similarity of handshape and motion. By mapping isolated dictionary exemplars into this space, the method enables direct, embedding‑based sign matching and can generalise to unseen signs using only dictionary examples. Experiments on ASL‑Citizen, ChaLearn OSLWL, and BOBSL CSLR2 benchmarks show strong cross‑dataset, cross‑task, and cross‑language performance, outperforming prior approaches on American, British, and Spanish sign languages without benchmark‑specific supervision.

By Ryan Wong, Youngjoon Jang, Liliane Momeni, G\"ul Varol, Andrew Zisserman