PHOSA presents a photorealistic 3D sign avatar modeling framework that addresses the need for accurate communication with the Deaf community. The authors introduce MVSign, a multi‑view Chinese sign language dataset co‑designed with Deaf experts, and a hybrid fitting pipeline for precise SMPL‑X annotation. Their decoupled avatar representation and motion‑aware sampling achieve high‑fidelity visual results on MVSign and generalize to monocular sign language videos.
arXiv:2605. 05367v2 Announce Type: replace-cross Abstract: Existing 3D sign language avatar reconstruction methods are developed and evaluated exclusively on Western sign languages, and no 3D parametric annotations exist for any Arabic Sign Language dataset, a gap that blocks the development of avatar-based accessibility applications for the Arab Deaf community.
By Eyad Alghamdi, Sattam Altuuaim, Obay Ghulam, Abdulrahman Qutah, Yousef Basoodan
arXiv:2605. 01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings.
By Sen Fang, Hongbin Zhong, Yanxin Zhang, Dimitris N. Metaxas
M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.
By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
SignGPT is a unified, pose‑based framework that performs gloss‑free sign language translation (SLT) and generation (SLG) by integrating part‑aware hierarchical representations of body, hand, and facial motion into a shared language model. It uses asymmetric multi‑token prediction and progressive training for bidirectional modeling, and is evaluated on How2Sign (ASL) and Phoenix‑2014T (DGS) with benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers demonstrates a sign‑to‑sign response pipeline, suggesting that unified modeling can support sign language conversation (SLC).
By Ronghui Li, Jun Dong, Zhongyuan Hu, Zunnan Xu, Jun Zhou, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Jie Guo, Xiu Li, Linchao Bao
arXiv:2609.14122v1 Announce Type: new
Abstract: We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer...
By Zhewen He (New York University Abu Dhabi), Junyi Yu (New York University Abu Dhabi), Haomian Huang (New York University Abu Dhabi), Zhenhua Li (ChatSign Technology), Yi Fang (New York University Abu Dhabi, ChatSign Technology)
arXiv:2606. 19352v1 Announce Type: cross Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities.
By Yiming Ni, Zhi-Qi Cheng, Jiayu Li, Wei Cheng
arXiv:2610.00881v1 Announce Type: new
Abstract: Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation...
By Ozge Mercanoglu Sincan, Anton Pelykh, Edward Fish, Harry Walsh, JianHe Low, Karahan Sahin, Oline Ranum, Sobhan Asasi, Steven Emery, Richard Bowden
SignRR is a new sign language production framework that combines retrieval of real sign motion segments with a learned refinement step to produce globally coherent signing sequences. It starts from a dictionary of authentic sign segments and refines them using a part-aware Residual VQ‑VAE, preserving fine hand articulation while handling temporal length differences in latent space. Experiments on PHOENIX14T and CSL‑Daily demonstrate state‑of‑the‑art back‑translation performance and competitive pose quality.
By Fidel Omar Tito Cruz, Angie Sanchez Marquina, Summy Farfan, Gissella Bejarano
arXiv:2204. 02803v2 Announce Type: replace-cross Abstract: Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D observations and because the signal is inherently spatiotemporal.
By Silvan Ferreira, Esdras Costa, Marcio Dahia, Jampierre Rocha
arXiv:2607. 03657v1 Announce Type: cross Abstract: Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues.
By Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain, Leonid Sigal, Hamzah Luqman
arXiv:2609.39273v1 Announce Type: new
Abstract: This report presents \textbf{MegaAvatar}, a controllable talking avatar generation framework built on top of the Wan2.2-TI2V-5B model. Compared with pr...
By Junyao Gao, Sibo Liu, Weidong Zhang, Cairong Zhao, Jun Zhang