ScreenSuite - The most comprehensive evaluation suite for GUI Agents!
Related stories
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
arXiv:2607. 25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction.
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
ERPBench introduces a new evaluation paradigm for computer-use agents that operate via screenshots and simulated actions, focusing on enterprise software such as ERP systems. The benchmark tests agents on a live, reproducible ERP platform and scores tasks against ground-truth database values, highlighting challenges like dense interfaces, multi-step interactions, and persistent record errors. Experiments with six agents show that strong general GUI performance does not translate to reliable enterprise outcomes, with many agents frequently saving incorrect data.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
arXiv:2606. 11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
arXiv:2605. 25160v2 Announce Type: replace Abstract: GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments.
Software Engineering for and with GUI Agent
arXiv:2608. 09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications.
Introducing AgentKit, new Evals, and RFT for agents
Today, we’re releasing new tools to help developers go from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents.
Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications
The study evaluates computer-use agents (CUAs) for blind users by conducting a three‑week diary study with eight participants using the OLLA prototype. Across 1,258 commands in 12 desktop applications, GPT‑5 achieved the highest success rate of 52.5%, while analysis uncovered failures in grounding, planning, constraint‑tracking, and termination. Interviews highlighted additional needs beyond automation for blind users.
Syll: Open-Source Personal Automation with Cross-Surface Execution
arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
arXiv:2606. 29705v1 Announce Type: new Abstract: Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models.
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
arXiv:2606. 03203v1 Announce Type: new Abstract: Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated.