arXiv Computer Vision By Md Tariquzzaman, Md Farhan Ishmam, Saiyma Sittul Muna, Md Kamrul Hasan, Hasan Mahmud

Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation

Read the original on arXiv Computer Vision →

The paper introduces BdSLIG, the first Bengali Sign Language Instruction Generation dataset, aimed at evaluating Vision Language Models on under-resourced SLIG tasks and long-tail visual concepts. It proposes Sign Parameter-Infused (SPI) prompting, which embeds standard sign parameters such as hand shape, motion, and orientation into textual prompts to improve zero-shot performance and produce more structured, reproducible instructions. The work seeks to promote inclusivity and advance sign language learning systems for under-resourced communities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 31

DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

DeicticVLA unifies three instruction modes—Language Instruction, Vision‑Language Instruction, and Visual Instruction—into a single text prompt and deictic mask framework, allowing a single pretrained Vision‑Language‑Action model to handle all modes. The approach uses text‑prompt completion and deictic gesture grounding, and evaluates various visual prompting methods and training strategies in simulation and real‑world tasks. Results show that two‑stage training improves deictic mask usage, and that Vision‑Language and Visual Instruction outperform Language Instruction on unseen expressions, appearance changes, and novel objects, achieving 100% success on unseen categories.

By Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii
arXiv Computation and Language
Sep 3

SignBind-LLM: Multi-Stage Modality Fusion for Sign Language Translation

SignBind-LLM introduces a modular framework for sign language translation that separates continuous signing, fingerspelling, and lipreading into dedicated expert streams. Each expert is pre‑trained independently on about two million pseudo‑gloss sequences, eliminating the need for manual gloss annotation. A lightweight transformer fuses the expert outputs, and a pre‑trained language model converts the fused pseudo‑glosses into fluent English, achieving state‑of‑the‑art performance on multiple benchmarks with lower training cost.

By Marshall Thomas, Edward Fish, Richard Bowden
arXiv AI
Aug 11

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

arXiv:2608. 09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress.

By Xiao Liu, Shiwei Gan, Yafeng Yin, Jiaxin Yin, Bowen Guo, Yaqi Sun, Zhiwei Jiang, Lei Xie
arXiv AI
Jul 7

ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

arXiv:2607. 03657v1 Announce Type: cross Abstract: Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues.

By Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain, Leonid Sigal, Hamzah Luqman