GUI-PRA: Process Reward Agent for GUI Tasks
arXiv:2509.23263v3 Announce Type: replace Abstract: Long-horizon GUI automation remains challenging due to error accumulation over extended interaction sequences. Process Reward Models (PRMs) provide...
The paper introduces Evidence-First Reflection (EFR), a two-stage reflector for desktop GUI agents that separates visual difference extraction from outcome verification. EFR uses Set-of-Marks annotations to locate action sites and candidate changed regions, filters action-relevant changes, and then makes a final judgment based on cleaned evidence. Experiments on OSWorld-Verified and WindowsAgentArena show that EFR improves reflector accuracy by 7.11% and increases end-to-end task success by roughly 5–6%.
arXiv:2509.23263v3 Announce Type: replace Abstract: Long-horizon GUI automation remains challenging due to error accumulation over extended interaction sequences. Process Reward Models (PRMs) provide...
arXiv:2606. 11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments.
arXiv:2610.01215v1 Announce Type: new Abstract: GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step...
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available tr...
arXiv:2607. 26041v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks.
arXiv:2608. 11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents.
arXiv:2607. 15193v1 Announce Type: new Abstract: Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent.
arXiv:2607. 25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction.
arXiv:2609.24094v1 Announce Type: cross Abstract: Visual analytics (VA) enables sensemaking through interactive visualization, but effective analysis often requires experts to translate high-level in...
arXiv:2608. 05587v1 Announce Type: new Abstract: Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution.
arXiv:2607. 17050v1 Announce Type: cross Abstract: GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery.
ComponentBench is a new benchmark that evaluates computer‑use agents at the component level on modern web UIs. It contains 97 canonical UI components and 2,910 programmatically verified tasks, along with cleaned human reference trajectories for measuring task success and interaction efficiency. The benchmark also offers a scalable pipeline for auditing structural difficulty and synthesizing failure analyses across tasks and component families.