arXiv Computer Vision

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

The paper introduces BdSLIG, the first Bengali Sign Language Instruction Generation dataset, aimed at evaluating Vision Language Models on under-resourced SLIG tasks and long-tail visual concepts. It proposes Sign Parameter-Infused (SPI) prompting, which embeds standard sign parameters such as hand shape, motion, and orientation into textual prompts to improve zero-shot performance and produce more structured, reproducible instructions. The work seeks to promote inclusivity and advance sign language learning systems for under-resourced communities.

arXiv Computer Vision
Aug 31

DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

DeicticVLA unifies three instruction modes—Language Instruction, Vision‑Language Instruction, and Visual Instruction—into a single text prompt and deictic mask framework, allowing a single pretrained Vision‑Language‑Action model to handle all modes. The approach uses text‑prompt completion and deictic gesture grounding, and evaluates various visual prompting methods and training strategies in simulation and real‑world tasks. Results show that two‑stage training improves deictic mask usage, and that Vision‑Language and Visual Instruction outperform Language Instruction on unseen expressions, appearance changes, and novel objects, achieving 100% success on unseen categories.

By Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii
arXiv Computation and Language
Sep 3

SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation

SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.

By Marshall Thomas, Edward Fish, Richard Bowden
arXiv AI
Aug 11

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

arXiv:2608. 09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress.

By Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie
arXiv AI
Jul 7

ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

arXiv:2607. 03657v1 Announce Type: cross Abstract: Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues.

By Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain, Leonid Sigal, Hamzah Luqman
arXiv Machine Learning
Jun 11

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.

By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv AI
Sep 4

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

The paper argues that BLEU-4, the prevailing metric for sign language translation (SLT), may not accurately reflect sign language proficiency because SLT models can exploit spurious correlations and spoken-language priors. By evaluating six SLT models on Phoenix-2014T and CSL-Daily, the authors show that higher BLEU-4 scores do not necessarily indicate better spatio-temporal understanding. They propose a new open-weight LLM QA protocol inspired by language-learning assessment, which better preserves salient content, aligns more closely with human rankings, and reveals differences between gloss-free and gloss-supervised systems that BLEU-4 obscures.

By Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden
arXiv Computer Vision
Sep 4

M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production

M3T introduces a discrete multi‑modal motion token system for sign language production, addressing the need for non‑manual features such as mouthings, eyebrow raises, gaze, and head movements. The approach couples FLAME’s expressive facial space with SMPL‑X body parameters and uses modality‑specific Finite Scalar Quantization VAEs to achieve high face codebook utilization (99.0%). Trained with an autoregressive transformer and a sign‑to‑text translation objective, M3T outperforms existing methods on three standard datasets, notably improving accuracy on NMFs‑CSL from 49.0% to 58.3% without large‑scale pre‑training.

By Alexandre Symeonidis-Herzig, Jianhe Low, Ozge Mercanoglu Sincan, Richard Bowden
arXiv Computer Vision
Aug 27

Recognising BSL Fingerspelling in Continuous Signing Sequences

The paper introduces FS23K, a large-scale British Sign Language fingerspelling dataset created through an iterative annotation framework. It also presents a recognition model that incorporates bi‑manual interactions and mouthing cues, achieving a halved character error rate compared to previous state‑of‑the‑art methods. These results underscore the dataset’s and model’s value for advancing sign language research and automated annotation pipelines.

By Alyssa Chan, Taein Kwon, Andrew Zisserman