SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
arXiv:2608. 09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks.
arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.
arXiv:2608. 09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks.
arXiv:2605. 01720v3 Announce Type: replace-cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in laboratory settings.
arXiv:2605. 31393v2 Announce Type: replace-cross Abstract: Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed vocabularies typical of real-world datasets.
SMART is a new framework that jointly tackles continuous sign language recognition (CSLR) and spotting by leveraging a multimodal large language model (MLLM) to generate motion descriptions as auxiliary semantic cues. It performs stable video‑text alignment with small batch sizes and introduces a Multi‑Scale Temporal Adapter to capture temporal interactions during transformer encoding. The framework also incorporates CSFormer, a CSLR‑guided spotting module that injects recognition‑derived gloss evidence into a boundary‑aware spotting network, enabling mutual benefit between recognition and spotting tasks.
arXiv:2607. 03657v1 Announce Type: cross Abstract: Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues.
Continuous sign language recognition (CSLR) aims to recognize gloss sequences from unsegmented sign videos under weak sequence-level supervision. However, existing methods rely on sentence-level gloss...
RAIDAL is an active learning framework for continuous sign language recognition that leverages the CTC decoder’s alignment peaks to focus sample selection on gloss‑aligned regions, thereby avoiding temporal redundancy in weakly aligned videos. By restricting representation‑based scoring to these decoder‑aligned gloss areas, RAIDAL improves data efficiency across multiple datasets and architectures, especially in large‑vocabulary, budget‑limited scenarios. The method requires no extra labeling cost and its implementation is publicly available on GitHub.
SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.
arXiv:2609.05742v1 Announce Type: cross Abstract: American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion represen...
arXiv:2603. 29219v2 Announce Type: replace-cross Abstract: Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community.
arXiv:2609.14122v1 Announce Type: new Abstract: We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer...
The paper introduces SignShift, a framework for visual-only sentence-level segmentation of continuous sign language videos. It uses a Temporal Difference Module that captures frame-to-frame feature variations across full-frame, facial, and hand cues, and a Segment Count Prediction module to guide boundary selection. Experiments on benchmark datasets show that SignShift outperforms existing methods, demonstrating its effectiveness for this challenging task.