arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.
By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
arXiv:2606. 05981v2 Announce Type: replace-cross Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1-step distilled student, the text encoder becomes the critical path.
By Yoshiyuki Ootani
arXiv:2606. 05981v1 Announce Type: cross Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1-step distilled student, the text encoder becomes the critical path.
By Yoshiyuki Ootani
arXiv:2608. 06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL).
By Saad Ahmed, Md Khalid Syfullaha
arXiv:2506. 10915v2 Announce Type: replace-cross Abstract: Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators.
By Jiancheng Huang, Gengwei Zhang, Zequn Jie, Siyu Jiao, Yinlong Qian, Ling Chen, Yunchao Wei, Lin Ma
Token-Budget Distillation (TBD) is a parameter‑efficient fine‑tuning framework that adapts video vision‑language models to a fixed token budget. It freezes the pretrained backbone, updates only LoRA adapters, and incorporates FlashVID visual token compression. TBD uses a dual‑path teacher‑student design with full‑token supervision and compressed student optimization, enabling the student to recover full‑token semantics while remaining efficient under aggressive token reduction.
By Xiaoyang Guo, Guoping Luo, Jusheng Zhang, Keze Wang, Wenhao Wang
arXiv:2607. 14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult.
By Wenxuan Chen, Wenjie Feng
Video DeltaNet (VDN) introduces a hybrid attention mechanism for livestream video generation, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). VDA updates memory once per frame, integrating spatial tokens, while separate output projections and learnable gates balance the two branches. Applied to MiniMax H3, VDN achieves a 14.5× speedup over the dense baseline, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs.
By Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng
RAIDAL is an active learning framework for continuous sign language recognition that leverages the CTC decoder’s alignment peaks to focus sample selection on gloss‑aligned regions, thereby avoiding temporal redundancy in weakly aligned videos. By restricting representation‑based scoring to these decoder‑aligned gloss areas, RAIDAL improves data efficiency across multiple datasets and architectures, especially in large‑vocabulary, budget‑limited scenarios. The method requires no extra labeling cost and its implementation is publicly available on GitHub.
By Rafael A. Diniz Augusto, Gabriel L. Oliveira, Erickson R. Nascimento
arXiv:2608. 13368v1 Announce Type: cross Abstract: This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Adversarial Network (GAN) to enhance communication for individuals with hearing impairments.
By Dingzhan Nong, Zhihao Ren, Ziqi Li, Tim Lo
Continuous sign language recognition (CSLR) is a key technology for accessibility, yet its development remains limited by the high cost of annotating continuous video streams. Active learning offers a...
The paper introduces AttnSign, a vision‑language model that improves sign language translation by steering spatial‑temporal attention. It first supervises attention on sign‑relevant regions such as faces and hands in each frame, then uses an RL‑based motion‑cadence method to focus on keyframes. Experiments on How2Sign and OpenASL show AttnSign consistently outperforms existing methods.
By Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao