Hugging Face Trending Papers

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

arXiv Computer Vision
Aug 27

SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting

SMART is a new framework that jointly tackles continuous sign language recognition (CSLR) and spotting by leveraging a multimodal large language model (MLLM) to generate motion descriptions as auxiliary semantic cues. It performs stable video‑text alignment with small batch sizes and introduces a Multi‑Scale Temporal Adapter to capture temporal interactions during transformer encoding. The framework also incorporates CSFormer, a CSLR‑guided spotting module that injects recognition‑derived gloss evidence into a boundary‑aware spotting network, enabling mutual benefit between recognition and spotting tasks.

By Eunjee Choi, JungHoon Sung, Seongwhan Cho, Chu Xin, Younggeun Choi
arXiv AI
Jul 7

ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

arXiv:2607. 03657v1 Announce Type: cross Abstract: Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues.

By Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain, Leonid Sigal, Hamzah Luqman
arXiv Machine Learning
Jun 11

Corpus Augmentation for Sign Language Translation via LLM-Guided Video Stitching

arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.

By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv Computer Vision
Sep 11

RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition

RAIDAL is an active learning framework for continuous sign language recognition that leverages the CTC decoder’s alignment peaks to focus sample selection on gloss‑aligned regions, thereby avoiding temporal redundancy in weakly aligned videos. By restricting representation‑based scoring to these decoder‑aligned gloss areas, RAIDAL improves data efficiency across multiple datasets and architectures, especially in large‑vocabulary, budget‑limited scenarios. The method requires no extra labeling cost and its implementation is publicly available on GitHub.

By Rafael A. Diniz Augusto, Gabriel L. Oliveira, Erickson R. Nascimento
arXiv Computer Vision
Sep 3

SignMatch: Matching Dictionary Signs to Continuous Sign Language Video

SignMatch introduces a prototype‑structured embedding space that learns to match dictionary sign videos with continuous sign language footage based solely on visual similarity of handshape and motion. By mapping isolated dictionary exemplars into this space, the method enables direct, embedding‑based sign matching and can generalise to unseen signs using only dictionary examples. Experiments on ASL‑Citizen, ChaLearn OSLWL, and BOBSL CSLR2 benchmarks show strong cross‑dataset, cross‑task, and cross‑language performance, outperforming prior approaches on American, British, and Spanish sign languages without benchmark‑specific supervision.

By Ryan Wong, Youngjoon Jang, Liliane Momeni, G\"ul Varol, Andrew Zisserman