arXiv Computer Vision By Bowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu, Kuizhuang Liu, Zhiwei Jiang, Lei Xie

Seeing Semantic Shift: Difference-Aware Sentence-Level Temporal Segmentation of Sign Language Videos

Read the original on arXiv Computer Vision →

The paper introduces SignShift, a framework for visual-only sentence-level segmentation of continuous sign language videos. It uses a Temporal Difference Module that captures frame-to-frame feature variations across full-frame, facial, and hand cues, and a Segment Count Prediction module to guide boundary selection. Experiments on benchmark datasets show that SignShift outperforms existing methods, demonstrating its effectiveness for this challenging task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

SMART is a new framework that jointly tackles continuous sign language recognition (CSLR) and spotting by leveraging a multimodal large language model (MLLM) to generate motion descriptions as auxiliary semantic cues. It performs stable video‑text alignment with small batch sizes and introduces a Multi‑Scale Temporal Adapter to capture temporal interactions during transformer encoding. The framework also incorporates CSFormer, a CSLR‑guided spotting module that injects recognition‑derived gloss evidence into a boundary‑aware spotting network, enabling mutual benefit between recognition and spotting tasks.

By Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin, Younggeun Choi
arXiv Computer Vision
Sep 18

Attention-Steered Vision-Language Models for Sign Language Translation

The paper introduces AttnSign, a vision‑language model that improves sign language translation by steering spatial‑temporal attention. It first supervises attention on sign‑relevant regions such as faces and hands in each frame, then uses an RL‑based motion‑cadence method to focus on keyframes. Experiments on How2Sign and OpenASL show AttnSign consistently outperforms existing methods.

By Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao
arXiv Machine Learning
Jun 11

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.

By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv AI
Aug 11

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

arXiv:2608. 09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress.

By Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie