ScreenSuite - The most comprehensive evaluation suite for GUI Agents!
Related stories
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
arXiv:2607. 25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction.
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
arXiv:2606. 11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
arXiv:2605. 25160v2 Announce Type: replace Abstract: GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments.
Software Engineering for and with GUI Agent
arXiv:2608. 09278v1 Announce Type: cross Abstract: GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications.
Introducing AgentKit, new Evals, and RFT for agents
Today, we’re releasing new tools to help developers go from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents.
Syll: Open-Source Personal Automation with Cross-Surface Execution
arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
arXiv:2606. 29705v1 Announce Type: new Abstract: Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models.
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
arXiv:2606. 03203v1 Announce Type: new Abstract: Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated.
Plover: Steering GUI Agents through Plan-Centric Interaction
arXiv:2607. 15193v1 Announce Type: new Abstract: Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and evolving interface states can cause autonomous agents to drift from user intent.
A History-Aware Visually Grounded Critic for Computer Use Agents
arXiv:2606. 11078v1 Announce Type: new Abstract: Various test-time interventions for Computer Use Agents (CUAs), including critic models, have been developed to improve performance through pre-execution action evaluation in complex Graphical User Interface (GUI) environments.