arXiv AI

Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards

arXiv:2606. 19352v1 Announce Type: cross Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities.

arXiv Computer Vision
2d ago

Machine Translation for Sign Languages

arXiv:2610.00881v1 Announce Type: new Abstract: Sign language machine translation has progressed substantially over the past decade, evolving from isolated sign recognition to end-to-end translation...

By Ozge Mercanoglu Sincan, Anton Pelykh, Edward Fish, Harry Walsh, JianHe Low, Karahan Sahin, Oline Ranum, Sobhan Asasi, Steven Emery, Richard Bowden
arXiv Computer Vision
Aug 27

Recognising BSL Fingerspelling in Continuous Signing Sequences

The paper introduces FS23K, a large-scale British Sign Language fingerspelling dataset created through an iterative annotation framework. It also presents a recognition model that incorporates bi‑manual interactions and mouthing cues, achieving a halved character error rate compared to previous state‑of‑the‑art methods. These results underscore the dataset’s and model’s value for advancing sign language research and automated annotation pipelines.

By Alyssa Chan, Taein Kwon, Andrew Zisserman
arXiv Computer Vision
Sep 3

SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

SignMatch introduces a prototype‑structured embedding space that learns to match dictionary sign videos with continuous sign language footage based solely on visual similarity of handshape and motion. By mapping isolated dictionary exemplars into this space, the method enables direct, embedding‑based sign matching and can generalise to unseen signs using only dictionary examples. Experiments on ASL‑Citizen, ChaLearn OSLWL, and BOBSL CSLR2 benchmarks show strong cross‑dataset, cross‑task, and cross‑language performance, outperforming prior approaches on American, British, and Spanish sign languages without benchmark‑specific supervision.

By Ryan Wong, Youngjoon Jang, Liliane Momeni, G\"ul Varol, Andrew Zisserman
arXiv Computer Vision
Sep 21

SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

SignGPT is a unified, pose‑based framework that performs gloss‑free sign language translation (SLT) and generation (SLG) by integrating part‑aware hierarchical representations of body, hand, and facial motion into a shared language model. It uses asymmetric multi‑token prediction and progressive training for bidirectional modeling, and is evaluated on How2Sign (ASL) and Phoenix‑2014T (DGS) with benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers demonstrates a sign‑to‑sign response pipeline, suggesting that unified modeling can support sign language conversation (SLC).

By Ronghui Li, Jun Dong, Zhongyuan Hu, Zunnan Xu, Jun Zhou, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Jie Guo, Xiu Li, Linchao Bao
arXiv Computer Vision
Sep 25

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

PHOSA introduces MVSign, the first multi‑view Chinese sign language dataset co‑designed with Deaf experts, featuring diverse gestures and rich annotations. The authors develop a hybrid fitting pipeline for accurate SMPL‑X annotation and propose a decoupled sign avatar representation that isolates body, head, and hand components, coupled with a motion‑aware sampling strategy to handle motion blur and balance gesture diversity. Experiments show high‑fidelity visual results on MVSign, especially in detailed hand and facial regions, and good generalization to in‑the‑wild monocular sign language videos.

By Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li
Hugging Face Trending Papers
Sep 24

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

PHOSA presents a photorealistic 3D sign avatar modeling framework that addresses the need for accurate communication with the Deaf community. The authors introduce MVSign, a multi‑view Chinese sign language dataset co‑designed with Deaf experts, and a hybrid fitting pipeline for precise SMPL‑X annotation. Their decoupled avatar representation and motion‑aware sampling achieve high‑fidelity visual results on MVSign and generalize to monocular sign language videos.