JarvisGUI is a new benchmark that tests GUI agents on cross-device workflows involving Android, Windows, and Ubuntu, requiring transfer of intermediate results and coordination across heterogeneous platforms. It formulates tasks as input-output transformations under a lightweight type system, enabling automatic composition of multi-step, cross-device workflows and dynamic evaluation within a unified framework. The benchmark reveals that state-of-the-art open-source GUI agents struggle with state-transfer awareness, cross-platform contextual reasoning, and long-horizon dependency management, exposing a critical capability gap invisible to existing benchmarks.
By Zixiang Chen, Yuheng Lu, Zihao Cheng, Zeming Liu, Jizeng Bai, Ziye Huang, Zhiyin Lin, Zihan Li, Yuhang Guo, Yunhong Wang, Haifeng Wang
arXiv:2607. 13027v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action.
By Hongru Cai, Yongqi Li, Ran Wei, Wenjie Li
arXiv:2606.20487v2 Announce Type: replace
Abstract: Computer use agents are expanding from single-device operation toward cross-device systems that coordinate tasks across heterogeneous environments....
By Shu Yao, Yuhua Luo, Qian Long, Jingru Fan, Yuheng Wang, Lin Wu, Zhuoyuan Yu, Yufan Dang, Huatao Li, Chen Qian
arXiv:2608.23035v1 Announce Type: new
Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capabili...
By Yi Zhu, Xiongwei Wu, Qiyi Wang, Tingyu Qu, Jiajun Liu, Sihan Cao, Long Chen, Weigao Sun, Feida Zhu, Yiran Zhong, Steven Hoi
MobileGym is a browser-hosted, lightweight simulation platform designed for mobile GUI agent research. It offers verifiable outcome signals via deterministic, JSON-based state judging and supports scalable online reinforcement learning with hundreds of parallel instances on a single server. The platform includes a declarative task-definition framework, a structured AnswerSheet protocol, and a benchmark of 416 parameterized tasks across 28 apps, demonstrating strong sim-to-real transfer in a case study.
By Dingbang Wu, Rui Hao, Haiyang Wang, Shuzhe Wu, Han Xiao, Zhenghong Li, Bojiang Zhou, Zheng Ju, Zichen Liu, Lue Fan, Zhaoxiang Zhang
arXiv:2608. 05729v1 Announce Type: new Abstract: As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time.
By Xinshuang Liu, Runfa Blark Li, Shaoxiu Wei, Xin Lin, Truong Nguyen