arXiv Computation and Language

SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation

SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.

arXiv Computer Vision
Aug 27

Recognising BSL Fingerspelling in Continuous Signing Sequences

The paper introduces FS23K, a large-scale British Sign Language fingerspelling dataset created through an iterative annotation framework. It also presents a recognition model that incorporates bi‑manual interactions and mouthing cues, achieving a halved character error rate compared to previous state‑of‑the‑art methods. These results underscore the dataset’s and model’s value for advancing sign language research and automated annotation pipelines.

By Alyssa Chan, Taein Kwon, Andrew Zisserman
arXiv Machine Learning
Jun 11

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.

By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv Computer Vision
Sep 4

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.

By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
arXiv Computer Vision
Sep 4

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek is a new method for learning transferable sign representations that enables efficient retrieval of signs from dictionaries using only a query video. It employs contrastive learning with saliency‑guided articulator masking, aligning same‑gloss signs across signers while focusing on the single most critical articulator per sign. Trained on 266K samples from multiple sign languages, SignSeek achieves state‑of‑the‑art cross‑corpus retrieval performance and zero‑shot generalisation to unseen British Sign Language, also improving isolated sign recognition and subtitle alignment.

By Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden
arXiv Computer Vision
Sep 11

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

The paper introduces BdSLIG, the first Bengali Sign Language Instruction Generation dataset, aimed at evaluating Vision Language Models on under-resourced SLIG tasks and long-tail visual concepts. It proposes Sign Parameter-Infused (SPI) prompting, which embeds standard sign parameters such as hand shape, motion, and orientation into textual prompts to improve zero-shot performance and produce more structured, reproducible instructions. The work seeks to promote inclusivity and advance sign language learning systems for under-resourced communities.

By Md Tariquzzaman, Md Farhan Ishmam, Saiyma Sittul Muna, Md Kamrul Hasan, Hasan Mahmud
arXiv AI
Jun 19

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

arXiv:2605. 31393v2 Announce Type: replace-cross Abstract: Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed vocabularies typical of real-world datasets.

By Pedro Dal Bianco, Jean Paul Nunes Reinhold, Oscar Stanchi, Facundo Quiroga, Franco Ronchetti, Ulisses Brisolara Corr\^ea
arXiv Computer Vision
Aug 26

Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition

The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.

By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin
arXiv AI
Aug 12

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

arXiv:2608. 10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited.

By Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta