arXiv Computer Vision
Sep 21

SignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation

SignGPT is a unified, pose‑based framework that performs gloss‑free sign language translation (SLT) and generation (SLG) by integrating part‑aware hierarchical representations of body, hand, and facial motion into a shared language model. It uses asymmetric multi‑token prediction and progressive training for bidirectional modeling, and is evaluated on How2Sign (ASL) and Phoenix‑2014T (DGS) with benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers demonstrates a sign‑to‑sign response pipeline, suggesting that unified modeling can support sign language conversation (SLC).

By Ronghui Li, Jun Dong, Zhongyuan Hu, Zunnan Xu, Jun Zhou, Liyuan Chen, Shuoling Liu, Jiangpeng Yan, Jie Guo, Xiu Li, Linchao Bao
arXiv Computer Vision
Aug 27

Recognising BSL Fingerspelling in Continuous Signing Sequences

The paper introduces FS23K, a large-scale British Sign Language fingerspelling dataset created through an iterative annotation framework. It also presents a recognition model that incorporates bi‑manual interactions and mouthing cues, achieving a halved character error rate compared to previous state‑of‑the‑art methods. These results underscore the dataset’s and model’s value for advancing sign language research and automated annotation pipelines.

By Alyssa Chan, Taein Kwon, Andrew Zisserman