Machine Translation for Sign Languages
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 19352v1 Announce Type: cross Abstract: Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities.
arXiv:2608. 06407v1 Announce Type: cross Abstract: Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem.
SignGPT is a unified, pose‑based framework that performs gloss‑free sign language translation (SLT) and generation (SLG) by integrating part‑aware hierarchical representations of body, hand, and facial motion into a shared language model. It uses asymmetric multi‑token prediction and progressive training for bidirectional modeling, and is evaluated on How2Sign (ASL) and Phoenix‑2014T (DGS) with benchmark comparisons, qualitative analyses, and component ablations. An exploratory study with 12 Deaf ASL signers demonstrates a sign‑to‑sign response pipeline, suggesting that unified modeling can support sign language conversation (SLC).
arXiv:2204. 02803v2 Announce Type: replace-cross Abstract: Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D observations and because the signal is inherently spatiotemporal.
The paper introduces FS23K, a large-scale British Sign Language fingerspelling dataset created through an iterative annotation framework. It also presents a recognition model that incorporates bi‑manual interactions and mouthing cues, achieving a halved character error rate compared to previous state‑of‑the‑art methods. These results underscore the dataset’s and model’s value for advancing sign language research and automated annotation pipelines.
arXiv:2605. 01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings.