arXiv Computer Vision

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

The paper surveys the evolution of smart glasses from simple capture devices to first‑person intelligence platforms that integrate human perception, context, and action. It introduces a unified framework that formalizes data flow, hardware capabilities, and seven foundational capabilities, and presents an L0‑L5 hierarchy for capture to embodied action. The study also maps nine application scenes, proposes a nine‑dimensional deployment framework, and outlines an evidence ladder for evaluation and trustworthiness.

arXiv Computer Vision
Sep 18

AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization

The paper surveys the evolution of smart glasses into AI smart glasses, framing them as wearable intelligence platforms that integrate egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world constraints for personalized assistance. It organizes the discussion into four dimensions: hardware foundations, wearable intelligence, interaction design, and application scenarios across healthcare, accessibility, learning, daily life, tourism, and industry. The authors identify five cross-cutting research challenges—next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models—to guide future work.

By Xu Yuan, Yi Wang, Zhuohang Jiang, Haohao Qu, Yujuan Ding, Shanru Lin, Guoliang Xing, Hongxia Yang, Jiannong Cao, Qing Li, Wenqi Fan
arXiv Computer Vision
3d ago

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

arXiv:2609.39378v1 Announce Type: new Abstract: Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking evolvin...

By Shulin Tian, Junsu Kim, Shuai Liu, Hao Li, Yujiao Shen, Sihan Li, Zhe Yang, Yeongon Kim, Feiyu Li, Jialin Wu, Yichi Zhang, Wenhui Wang, Runmao Yao, Yuhao Dong, Zhaoxi Chen, Fangzhou Hong, Antonino Furnari, Jingkang Yang, Hongyuan Zhu, Ziwei Liu
Hugging Face Trending Papers
Jul 13

From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence

Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-action policies, and world models have advanced this goal, while World Action Models (WAMs) are particularly promising because they connect candidate interventions with predicted consequences.

arXiv AI
Sep 16

World Models for Embodied Intelligence: From Plausible to Controllable to Actionable

arXiv:2609.16697v1 Announce Type: cross Abstract: World models connect perception and decision-making in embodied intelligence by maintaining hidden state, anticipating consequences, comparing interv...

By Nanjie Yao, Hao Wang, Chong Cheng, Zhikang Chen, Wenzhe Li, Jiafei Lyu, Li Shen, Peilin Zhao, Zongqing Lu, Gao Huang, Steven Hoi, Dacheng Tao, Deheng Ye
arXiv AI
Jun 30

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

arXiv:2511. 17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday human-built environments.

By Juntao Cheng, Wanyue Zhang, Zhiwei Yu, Shuo Ren, Zheqi He, Shaoxuan Xie, Guocai Yao, Jieru Lin, B\"orje F. Karlsson, Jiajun Zhang
arXiv AI
Jul 29

ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop

arXiv:2607. 24770v1 Announce Type: new Abstract: Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions, track task progress, reason about spatial state, and recover from errors while performing physical actions.

By Azizul Zahid, Subrata Biswas, Bashima Islam, Sai Swaminathan
arXiv AI
Jul 7

Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems

arXiv:2607. 03283v1 Announce Type: new Abstract: Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal observations, robot states, human demonstrations, and task contexts into structured representations, decisions, trajectories, control references, and system services.

By Junwu Xiong, Jiaxuan Gao, Wei Chai, Renxing Chen, Yuzhen Li, Yu Guo, Yucheng Guo, Mingxi Luo, Wenyang Ma, Yiyun Mou, Yifei Zhang, Chen Zhou, Yongjian Guo
arXiv Computer Vision
2d ago

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

PhysVista is a new benchmark that evaluates physical intelligence in Vision‑Language Models (VLMs) by integrating perception, reasoning, and plausibility assessment into a closed cognitive loop. It distinguishes between event‑level and scale‑level reasoning and tests models on both real‑world and AI‑generated videos to provide a holistic, fine‑grained analysis of physical understanding. Experiments show significant gaps in VLMs’ physical reasoning and plausibility assessment, underscoring the need for more principled, physically grounded multimodal designs.

By Xinge Peng, Yiting Lu, Tianwu Zhi, Wen Wen, Jianzhao Liu, Xin Li, Zhibo Chen