Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often...
arXiv:2609.38008v1 Announce Type: new
Abstract: Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical us...
By Tongbo Chen, Junbo Niu, Zhengxi Lu, Niu Lian, Fei Tang, Yuchen Yan, Yike Hong, Yong Du, Yizhou Liu, Bofan Chen, Yongliang Shen
arXiv:2607. 13027v1 Announce Type: cross Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action.
By Hongru Cai, Yongqi Li, Ran Wei, Wenjie Li
AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.
By Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang
The paper introduces EvoSkill-GUI, a training‑free framework that enables GUI agents to evolve their skills during deployment. Each skill is packaged with metadata, executable plans, and recovery rules, and the system follows a reflect‑revise‑reuse loop where the agent instantly revises skills based on execution feedback. Experiments on MobileWorld, AndroidWorld, and OSWorld show consistent performance gains up to +16.2% without any additional training.
By Bofan Chen, Boxuan Zhang, Fei Tang, Zhengxi Lu, Yong Du, Tongbo Chen, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
arXiv:2608. 15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution.
By Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng