arXiv AI By Mengzhang Li, Yuan Yao

OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

Read the original on arXiv AI →

arXiv:2607. 03213v1 Announce Type: cross Abstract: We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary focus on blind and low-vision users.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 3

VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval

arXiv:2607. 02371v1 Announce Type: cross Abstract: Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belongings, recognizing familiar faces, or handling cash remain persistent obstacles to personal autonomy.

By Cristian-Gabriel Florea, Stelian Sp\^inu
arXiv Computer Vision
Aug 26

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

The paper surveys the evolution of smart glasses from simple capture devices to first‑person intelligence platforms that integrate human perception, context, and action. It introduces a unified framework that formalizes data flow, hardware capabilities, and seven foundational capabilities, and presents an L0‑L5 hierarchy for capture to embodied action. The study also maps nine application scenes, proposes a nine‑dimensional deployment framework, and outlines an evidence ladder for evaluation and trustworthiness.

By Jiangning Zhang, Haojun Chen, Yong Liu
arXiv Computer Vision
Sep 18

AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization

The paper surveys the evolution of smart glasses into AI smart glasses, framing them as wearable intelligence platforms that integrate egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world constraints for personalized assistance. It organizes the discussion into four dimensions: hardware foundations, wearable intelligence, interaction design, and application scenarios across healthcare, accessibility, learning, daily life, tourism, and industry. The authors identify five cross-cutting research challenges—next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models—to guide future work.

By Xu Yuan, Yi Wang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Shanru Lin, Guoliang Xing, Hongxia Yang, Jiannong Cao, Qing Li, Wenqi Fan
arXiv Machine Learning
Sep 25

Edge AI on Constrained Devices for Binary Sleep-Wake Classification in Dynamic Environments

This paper introduces an Edge AI system that classifies sleep and wake states on constrained devices using a multimodal pipeline on an ESP32‑S3 microcontroller. It fuses inertial head‑movement sensing with visual pose classification, running in parallel under FreeRTOS to meet real‑time constraints. The two‑stage detection achieves 96.5 % accuracy for motion‑based detection and 89 % for pose classification, proving robust binary sleep‑wake classification in mobile scenarios.

By Stefan Reitmann, Lena Oden