arXiv AI

Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model

arXiv:2608. 06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL).

arXiv Computation and Language
Sep 3

SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation

SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.

By Marshall Thomas, Edward Fish, Richard Bowden
arXiv AI
Aug 12

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

arXiv:2608. 10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited.

By Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
arXiv Computer Vision
Sep 4

SignSeek: Learning Transferable Representations for Sign Dictionary Retrieval

SignSeek is a new method for learning transferable sign representations that enables efficient retrieval of signs from dictionaries using only a query video. It employs contrastive learning with saliency‑guided articulator masking, aligning same‑gloss signs across signers while focusing on the single most critical articulator per sign. Trained on 266K samples from multiple sign languages, SignSeek achieves state‑of‑the‑art cross‑corpus retrieval performance and zero‑shot generalisation to unseen British Sign Language, also improving isolated sign recognition and subtitle alignment.

By Sobhan Asasi, Ozge Mercanoglu Sincan, Richard Bowden
arXiv Computer Vision
Aug 26

Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition

The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.

By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin
arXiv AI
Sep 4

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

The paper argues that BLEU-4, the prevailing metric for sign language translation (SLT), may not accurately reflect sign language proficiency because SLT models can exploit spurious correlations and spoken-language priors. By evaluating six SLT models on Phoenix-2014T and CSL-Daily, the authors show that higher BLEU-4 scores do not necessarily indicate better spatio-temporal understanding. They propose a new open-weight LLM QA protocol inspired by language-learning assessment, which better preserves salient content, aligns more closely with human rankings, and reveals differences between gloss-free and gloss-supervised systems that BLEU-4 obscures.

By Oline Ranum, Edward Fish, Simon Hadfield, Richard Bowden
arXiv Computation and Language
Sep 25

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

The paper introduces Smol‑VL‑BLV, a compact vision‑language model designed for blind and low‑vision users. It employs a 500M decoder transformer with teacher‑student distillation and Group Relative Policy Optimization to add spatial detail, directional cues, and hazard detection to post‑training. After a lightweight finetuning step, the model achieves significant gains on spatial, social, OCR, and VQA benchmarks while remaining under 450 MB and running entirely offline on a mid‑range Android phone.

By Rishabh Choudhary, Shreyansh Raj, Umesh Goyal, Shubh Kashyap, Shrestha Kumar, Sushovan Jena, Komal Kumar, Hisham Cholakkal, Aditya Nigam