arXiv AI By Yudong Zhang (Honor Device Co., Ltd), Lei Hu (Honor Device Co., Ltd), Daoyang Liu (The Chinese University of Hong Kong, Hong Kong, China), Jiawei Liu (Honor Device Co., Ltd), Yangfan Luo (Honor Device Co., Ltd), Xingyu Liu (Honor Device Co., Ltd), Zuojian Wang (Honor Device Co., Ltd), Zhilin Gao (Honor Device Co., Ltd)

Teach-and-Repeat: Accurately Extracting Operational Knowledge from Mobile Screen Demonstrations to Empower GUI Agents

Read the original on arXiv AI →

arXiv:2606. 12817v1 Announce Type: new Abstract: Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 15

GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models

arXiv:2606. 12817v2 Announce Type: replace Abstract: Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension.

By Yudong Zhang (Honor Device Co., Ltd), Lei Hu (Honor Device Co., Ltd), Daoyang Liu (The Chinese University of Hong Kong, Hong Kong, China), Jiawei Liu (Honor Device Co., Ltd), Yangfan Luo (Honor Device Co., Ltd), Zhilin Gao (Honor Device Co., Ltd), Zuojian Wang (Honor Device Co., Ltd)
arXiv AI
Jun 11

Grounding Computer Use Agents on Human Demonstrations

arXiv:2511. 07332v2 Announce Type: replace-cross Abstract: Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements.

By Aarash Feizi, Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Kaixin Li, Rabiul Awal, Xing Han L\`u, Johan Obando-Ceron, Juan A. Rodriguez, Nicolas Chapados, David Vazquez, Adriana Romero-Soriano, Reihaneh Rabbany, Perouz Taslakian, Christopher Pal, Spandana Gella, Sai Rajeswar
arXiv AI
Sep 2

Towards Generalizable Visually Grounded Exploration of Household Devices

The paper introduces VGEBench, a new benchmark for evaluating Vision‑Language Models (VLMs) on generalizable, visually grounded exploration of household devices. Unlike existing datasets that rely on static images or annotated trajectories, VGEBench employs a logic‑driven state machine to simulate multi‑turn interaction loops, requiring agents to actively perceive, act, and refine their actions to achieve goals. Experiments show that current VLMs struggle to translate semantic knowledge into physical execution and to maintain long‑horizon state tracking.

By Linhao Zheng, Zeming Liu, Wangke Chen, Li Zeng, Wanxiang Che, Heyan Huang, Yuhang Guo
arXiv AI
Aug 24

Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation

The paper introduces CRATE, a two‑stage vision‑language model framework that evaluates mobile agents by reasoning about each step’s consequences and aggregating this evidence to assess task completion. It also presents CRATE‑S, an extension that evaluates operational safety. Experiments show CRATE and CRATE‑S outperform existing benchmarks, achieving high F1‑scores on AndroidWorld and MobileRisk datasets.

By Pengshuai Yang, Zijing Gao, Xue Yu, Benhui Zhuang, Bo Yuan, Junlan Feng