arXiv Computation and Language By Bizhe Bai, Jiakang Yuan, Hongming Wu, Xinyue Wang, Jie Ren, Siyao Chen, Yuchen Ya, Fan Bai, Pai Peng, Huafeng Qin, Tao Chen

Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

Read the original on arXiv Computation and Language →

The paper surveys efficient GUI agents, emphasizing that practical deployment requires more than task success—it must also minimize context, computation, action budget, and runtime overhead. It reviews key efficiency dimensions—observation, memory, action, and planner/system—and identifies recurring strategies such as selective reading, global-to-local visual allocation, recoverable memory, verification-aware control, and hybrid runtimes. The authors highlight open challenges, including accurate verifier cost accounting, benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 16

EchoPath: Execution-Level Replayable Memory for GUI Agents

EchoPath is a model‑agnostic framework that transforms validated GUI interaction trajectories into standardized, parameter‑controlled memories, enabling agents to replay specific GUI actions deterministically. Each memory records task intent, preconditions, input parameters, GUI evidence, validation provenance, and lifecycle state, and uses an image‑based target‑reaiming algorithm to adjust coordinates for the current screen before execution. Experiments on real computer‑use tasks show that EchoPath cuts median token cost by over 90 % and median execution time by about 60 %.

By Yao Zhao, Aditya Shanmugham, Swastik Roy, Yanxun Xu
arXiv AI
Sep 11

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

JarvisGUI is a new benchmark that tests GUI agents on cross-device workflows involving Android, Windows, and Ubuntu, requiring transfer of intermediate results and coordination across heterogeneous platforms. It formulates tasks as input-output transformations under a lightweight type system, enabling automatic composition of multi-step, cross-device workflows and dynamic evaluation within a unified framework. The benchmark reveals that state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a critical capability gap invisible to existing benchmarks.

By Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang, Zhiyin Lin, Zihan Li, Yuhang Guo, Yunhong Wang, Haifeng Wang