arXiv:2606.21811v2 Announce Type: replace-cross
Abstract: Coding tasks are typically complicated and require multiple capabilities, ranging from high-level planning to low-level implementation. While...
By Shubham Gandhi, Yiqing Xie, Atharva Naik, Ruichen Zhu, Carolyn Rose
arXiv:2606. 11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments.
By Jaewoo Lee, Zaid Khan, Archiki Prasad, Justin Chih-Yao Chen, Supriyo Chakraborty, Kartik Balasubramaniam, Sambit Sahu, Elias Stengel-Eskin, Hyunji Lee, Mohit Bansal
arXiv:2607. 03441v1 Announce Type: cross Abstract: LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked.
By Yanbo Wang, Jinhua Hao, Yuze Shi, Kun Yuan, Ming Sun
arXiv:2606. 31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant attention for their utility and versatility.
By Xueqiao Sun, Xiaohan Wang, Ludwig Schmidt, Serena Yeung-Levy, Yuhui Zhang
Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet they remain prone t...
AgentHorizon is a benchmark comprising 1,373 computer‑use tasks with instruction‑trajectory pairs collected from 166 hours of human‑recorded activity across three operating systems. The benchmark uses a paired design that swaps instructions to create negative tasks, testing judges’ ability to distinguish successful trajectories from those that completed a similar but incompatible request. Eleven judges were evaluated, with GPT‑5.5 achieving the highest balanced accuracy of 80.9% on the full trajectory subset, while tool‑use benefits some models but harms others, revealing large variability in judges’ performance.
By Xing Han L\`u, Dheeraj Vattikonda, Sina Hajimiri, Fatemeh Pesaran Zadeh, Parishad BehnamGhader, Ghazwa Darwiche, Amirhossein Kazemnejad, Christopher Pal, Alexandre Drouin, Siva Reddy