arXiv AI

Reflection with Action-Induced Visual Differences for Desktop GUI Agents

The paper introduces Evidence-First Reflection (EFR), a two-stage reflector for desktop GUI agents that separates visual difference extraction from outcome verification. EFR uses Set-of-Marks annotations to locate action sites and candidate changed regions, filters action-relevant changes, and then makes a final judgment based on cleaned evidence. Experiments on OSWorld-Verified and WindowsAgentArena show that EFR improves reflector accuracy by 7.11% and increases end-to-end task success by roughly 5–6%.

arXiv AI
Sep 1

GUI-PRA: Process Reward Agent for GUI Tasks

arXiv:2509.23263v3 Announce Type: replace Abstract: Long-horizon GUI automation remains challenging due to error accumulation over extended interaction sequences. Process Reward Models (PRMs) provide...

By Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, Shengyu Zhang
arXiv AI
Jun 10

A History-Aware Visually Grounded Critic for Computer Use Agents

arXiv:2606. 11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments.

By Jaewoo Lee, Zaid Khan, Archiki Prasad, Justin Chih-Yao Chen, Supriyo Chakraborty, Kartik Balasubramaniam, Sambit Sahu, Elias Stengel-Eskin, Hyunji Lee, Mohit Bansal
arXiv Computer Vision
2d ago

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

arXiv:2610.01215v1 Announce Type: new Abstract: GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step...

By Cheng Yang, Yifan Wu, Yutao Huang, Zhaohua Zhang, Beiduo Chen, Muxi Chen, Chenchen Zhao, Hexuan Deng, Haolin Yang, Geyuan Zhu, Sa Zhu, Jianhuan Zhuo, Qiuyong Xiao, Jianhao Ruan, Yiran Peng, Jiayi Zhang, Tian Ye, Xinlei Yu, Tianwen Jiang, Jihong Zhang, Yuyu Luo
arXiv AI
Aug 20

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

ComponentBench is a new benchmark that evaluates computer‑use agents at the component level on modern web UIs. It contains 97 canonical UI components and 2,910 programmatically verified tasks, along with cleaned human reference trajectories for measuring task success and interaction efficiency. The benchmark also offers a scalable pipeline for auditing structural difficulty and synthesizing failure analyses across tasks and component families.

By Tianchen Guan, Xinlei Lin, Royce Cheng-Yue, Xiangjun Wang, Shuyan Zhou