arXiv AI By Xiaokai Yan, Jingtao Ding, Yong Li, Zhiwen Yu

Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions

Read the original on arXiv AI →

arXiv:2608. 14132v1 Announce Type: cross Abstract: Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.

By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
arXiv AI
Sep 4

Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation

The paper presents a unified framework for proactive service agents, defining proactivity as an agent’s ability to infer service opportunities from incomplete signals and decide whether to remain silent, ask, assist, or act. It models this as a partially observable sequential decision process constrained by authorization and risk, integrating timing, content, and delivery into a single structured action. The authors categorize existing methods along a decision pipeline—state and need estimation, intervention gating, action construction, and feedback adaptation—and propose standardized metrics for evaluating triggering, timing, calibration, user burden, safety, and policy value across diverse interaction modalities.

By Yan Tang, Tingyu Cao, Yuanbo Tang, Huaze Tang, Keer Hu
arXiv Computer Vision
Aug 31

Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

The paper introduces GMA, a new benchmark for evaluating general mobile assistants in realistic, challenging scenarios. GMA expands on existing benchmarks by offering seven open‑source applications across diverse domains and 300 tasks organized into four difficulty tiers, ranging from simple actions to complex multi‑step workflows. The authors evaluate eight state‑of‑the‑art models, showing that performance drops sharply with task complexity, and conduct ablation studies on harness design—such as context retention and state tracking—to demonstrate how these choices can improve outcomes, especially for demanding workflows.

By Yiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang Liu
arXiv AI
Sep 2

UI-Venus-2 Technical Report

UI‑Venus‑2 is a general‑purpose foundation GUI agent that operates across mobile, web, and desktop environments using a unified closed‑loop reasoning‑action framework. The report details how the system expands environment coverage to over 170 multilingual mobile apps and native desktop OSes, scales task generation through a deep‑research pipeline, and enhances verification with trace‑level and sample‑level evaluators that use visual keypoints and multi‑model voting. Safety‑aware mechanisms are also incorporated to control consequential actions, positioning UI‑Venus‑2 as an efficient, open‑source tool for more generalizable, verifiable, and self‑reflective agents in real‑world applications.

By Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
arXiv Computation and Language
Aug 28

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

The paper introduces INTENT-AS-A-TOOL, a method that equips large language models with intent-targeted tools to provide a fine-grained, judge‑free signal of their commitment to specific behaviors during reasoning. By monitoring the probability of calling these intent tools, the authors can track how intent evolves throughout generation, complementing chain‑of‑thought monitoring and expanding post‑hoc labels into dense trajectories. The approach identifies critical steps for online intervention, demonstrating that action preferences are useful for detecting agentic misalignment in autonomous agents.

By Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu
arXiv AI
Jul 1

Xiaomi-GUI-0 Technical Report

arXiv:2606. 31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation.

By Wanxia Cao, Chengzhen Duan, Pei Fu, Pengzhi Gao, Niu Lian, Fazhan Liu, Hui Liu, Heng Qu, Qinzhuo Wu, Zhehao Yu, Tongbo Chen, Shiqi Cui, Anan Du, Shukai Jia, Yuanfa Li, Yike Liu, Wenchao Lu, Haoyuan Sun, Jiatong Sun, Cheng Tan, Yajie Wang, Changqiao Wu, Tao Xiong, Jiahui Yang, Yuxuan Yuan, Ruoceng Zhang, Shaojie Zhang, Jian Zhu, Jian Luan, Cong Zou