arXiv AI By Keren Artiaga (Victor), Yang Li (Victor), Ercan Engin Kuruoglu (Victor), Wai Kin (Victor), Chan

Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment

Read the original on arXiv AI →

arXiv:2608. 16804v1 Announce Type: new Abstract: Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100 distinct sign languages are severely lacking.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

SMART is a new framework that jointly tackles continuous sign language recognition (CSLR) and spotting by leveraging a multimodal large language model (MLLM) to generate motion descriptions as auxiliary semantic cues. It performs stable video‑text alignment with small batch sizes and introduces a Multi‑Scale Temporal Adapter to capture temporal interactions during transformer encoding. The framework also incorporates CSFormer, a CSLR‑guided spotting module that injects recognition‑derived gloss evidence into a boundary‑aware spotting network, enabling mutual benefit between recognition and spotting tasks.

By Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin, Younggeun Choi
arXiv Computer Vision
Sep 18

Attention-Steered Vision-Language Models for Sign Language Translation

The paper introduces AttnSign, a vision‑language model that improves sign language translation by steering spatial‑temporal attention. It first supervises attention on sign‑relevant regions such as faces and hands in each frame, then uses an RL‑based motion‑cadence method to focus on keyframes. Experiments on How2Sign and OpenASL show AttnSign consistently outperforms existing methods.

By Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao
arXiv Computer Vision
Aug 26

Stack Transformer Based Spatial-Temporal Attention Model for Dynamic Sign Language and Fingerspelling Recognition

The paper introduces the Sequential Spatio-Temporal Attention Network (SSTAN), a Transformer-based architecture that replaces traditional Graph Convolutional Networks for sign language recognition. SSTAN uses a hierarchical, stacked design with Spatial Multi-Head Attention to model joint relationships within frames and Temporal Multi-Head Attention to capture long-range dependencies across frames, eliminating the need for predefined skeletal graphs. Experiments on large-scale datasets (WLASL, JSL, KSL) show that SSTAN, trained from scratch, achieves state‑of‑the‑art performance in fingerspelling categories and outperforms other skeleton‑only methods on WLASL, highlighting its data efficiency and ability to learn complex spatio‑temporal patterns.

By Koki Hirooka, Abu Saleh Musa Miah, Tatsuya Murakami, Md. Al Mehedi Hasan, Yong Seok Hwang, Jungpil Shin