arXiv:2606. 09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools.
By Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan
ComponentBench is a new benchmark that evaluates computer‑use agents at the component level on modern web UIs. It contains 97 canonical UI components and 2,910 programmatically verified tasks, along with cleaned human reference trajectories for measuring task success and interaction efficiency. The benchmark also offers a scalable pipeline for auditing structural difficulty and synthesizing failure analyses across tasks and component families.
By Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou
arXiv:2607. 22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs).
By Zedong Yu, Qianxing Li, Zhi Gao, Liuyu Xiang, Chenrui Shi, Yang Liu, Huiming Wu, Yujie Wei, Yuhao Fei, Yubo Fu, Zhaofeng He
arXiv:2606. 24551v1 Announce Type: new Abstract: Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and permitted actions.
By Xiao Zhou, Siyue Zhang, Yilun Zhao, Jinbiao Wei, Tingyu Song, Arman Cohan, Chen Zhao
arXiv:2610.00948v1 Announce Type: cross
Abstract: The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termi...
By Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai
arXiv:2607. 11818v1 Announce Type: cross Abstract: We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents.
By Kaixin Ma, Di Feng, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan