The paper introduces Smol‑VL‑BLV, a compact vision‑language model designed for blind and low‑vision users. It employs a 500M decoder transformer with teacher‑student distillation and Group Relative Policy Optimization to add spatial detail, directional cues, and hazard detection to post‑training. After a lightweight finetuning step, the model achieves significant gains on spatial, social, OCR, and VQA benchmarks while remaining under 450 MB and running entirely offline on a mid‑range Android phone.
By Rishabh Choudhary, Shreyansh Raj, Umesh Goyal, Shubh Kashyap, Shrestha Kumar, Sushovan Jena, Komal Kumar, Hisham Cholakkal, Aditya Nigam
arXiv:2607. 03213v1 Announce Type: cross Abstract: We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind and low-vision users.
By Mengzhang Li, Yuan Yao
arXiv:2607. 29353v1 Announce Type: cross Abstract: With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.
By Douwe den Blanken, Martin Lefebvre, Charlotte Frenkel
arXiv:2607. 03089v1 Announce Type: cross Abstract: HAR is increasingly expected to run continuously on edge devices, yet recent LLM-based methods remain hard to deploy: raw sensor prompts are long, cloud inference adds latency and privacy risk, and fine-tuned LLM pipelines turn general-purpose models into task-specific classifiers.
By Nirhoshan Sivaroopan, Albert Zomaya, Kanchana Thilakarathna
arXiv:2608. 15614v1 Announce Type: cross Abstract: The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget.
By Matteo Stoiber, Niels Buus Lassen
The paper presents a comprehensive benchmark for Domain Generalization (DG) in smartphone-based Human Activity Recognition (HAR), running over 410,000 experiments across multiple architectures, training objectives, initialization strategies, and architectural tweaks. It finds that individual DG components offer limited, highly conditional improvements, while combined configurations often yield stronger, sometimes super‑additive gains that depend on the model and shift scenario. The study also highlights that current source‑validation selection captures only a fraction of the potential oracle performance, underscoring the need for joint DG design and robust model‑selection methods.
By Ot\'avio Oliveira Napoli, Edson Borin
arXiv:2607. 26631v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments.
By Hansi Karunarathna, Nirhoshan Sivaroopan, Chamara Madarasingha, Anura Jayasumana, Kanchana Thilakarathna
arXiv:2606. 03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware.
By Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, Muhammet Esat Kalfaoglu
arXiv:2609.22277v1 Announce Type: cross
Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect elevated hazards or provide semantic environm...
By Ali Akarma, Adeel Ahmad, Toqeer Ali Syed
arXiv:2608. 06252v1 Announce Type: cross Abstract: Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL).
By Saad Ahmed, Md Khalid Syfullaha
Human Activity Recognition (HAR) from wearable sensors supports applications in healthcare, rehabilitation, fitness tracking, and smart environments. Yet, existing deep learning approaches require dataset-specific training, large labeled corpora, and repeated adaptation to new sensor settings or activity taxonomies.
arXiv:2607. 28627v1 Announce Type: cross Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
By Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem