arXiv Computation and Language

Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

The paper surveys efficient GUI agents, emphasizing that practical deployment requires more than task success—it must also minimize context, computation, action budget, and runtime overhead. It reviews key efficiency dimensions—observation, memory, action, and planner/system—and identifies recurring strategies such as selective reading, global-to-local visual allocation, recoverable memory, verification-aware control, and hybrid runtimes. The authors highlight open challenges, including accurate verifier cost accounting, benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.

arXiv AI
Sep 16

EchoPath: Execution-Level Replayable Memory for GUI Agents

EchoPath is a model‑agnostic framework that transforms validated GUI interaction trajectories into standardized, parameter‑controlled memories, enabling agents to replay specific GUI actions deterministically. Each memory records task intent, preconditions, input parameters, GUI evidence, validation provenance, and lifecycle state, and uses an image‑based target‑reaiming algorithm to adjust coordinates for the current screen before execution. Experiments on real computer‑use tasks show that EchoPath cuts median token cost by over 90 % and median execution time by about 60 %.

By Yao Zhao, Aditya Shanmugham, Swastik Roy, Yanxun Xu
arXiv AI
Sep 11

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

JarvisGUI is a new benchmark that tests GUI agents on cross-device workflows involving Android, Windows, and Ubuntu, requiring transfer of intermediate results and coordination across heterogeneous platforms. It formulates tasks as input-output transformations under a lightweight type system, enabling automatic composition of multi-step, cross-device workflows and dynamic evaluation within a unified framework. The benchmark reveals that state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a critical capability gap invisible to existing benchmarks.

By Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang, Zhiyin Lin, Zihan Li, Yuhang Guo, Yunhong Wang, Haifeng Wang
arXiv AI
Sep 2

GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments

GUI-CC is a benchmark designed to assess the contextual consistency of GUI world models when used as agent environments, rather than just one‑step next‑screen predictors. It includes two tracks: an offline reference‑action track that rolls models along real mobile GUI trajectories, and an online agent‑loop track where fixed probing agents interact with model‑generated UIs. The benchmark evaluates transition fidelity, plausibility, contextual consistency, and task progress across 500 offline trajectory tasks and 200 online tasks spanning 30 mobile apps.

By Lin Fu, Zheyuan Yang, Tianhui Zhang, Jinbiao Wei, Guo Gan, Boxu Liu, Yilun Zhao, Yu Rong
arXiv AI
Aug 18

From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems

arXiv:2608. 15127v1 Announce Type: cross Abstract: Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, environments, and persistent state.

By Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An, Tianyu Feng, Hanfeng Lu, Sheng Yao, Pu Guo, Yinghao Yu, Yizhou Shan, Bo Li, Binhang Yuan, Wei Wang
arXiv Computation and Language
Sep 21

MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks

MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.

By Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, Alex Pentland
arXiv AI
Sep 10

APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents

APPSim-Bench is a new benchmark for mobile GUI agents that uses controllable simulated apps to balance realism and reproducibility. It includes 557 tasks across 17 popular Chinese and English apps, with a coding-agent-assisted and human-verified workflow that ensures deterministic evaluation. Evaluation of 19 agents shows that autonomous mobile execution is still far from perfect, with the best model completing only 50.27% of tasks and many failures in longer workflows and numerical reasoning.

By Jintian Feng, Long Chen, Xiao Yu, Jiayi Dai, Chenglong Liu, Haoru Wang, Zizhen Xue, Yuxuan Shi, Ziyang Wang, Yichen Gong