Smol2Operator: Post-Training GUI Agents for Computer Use
Related stories
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
arXiv:2607. 25904v1 Announce Type: new Abstract: Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction.
HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical user interface (GUI) interactions, which are often...
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
arXiv:2607. 29320v1 Announce Type: new Abstract: Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments.
Syll: Open-Source Personal Automation with Cross-Surface Execution
arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
arXiv:2606. 29705v1 Announce Type: new Abstract: Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models.
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
arXiv:2607. 04425v2 Announce Type: replace-cross Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction.
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
arXiv:2605. 25160v2 Announce Type: replace Abstract: GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic environments.
HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents
arXiv:2609.38008v1 Announce Type: new Abstract: Computer use agents (CUAs) have demonstrated strong capabilities in completing digital tasks. However, existing CUAs either rely solely on graphical us...
Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents
The paper introduces GUI‑SD‑v2, an on‑policy self‑distillation framework that extends previous methods from GUI grounding to multi‑turn GUI interaction. It employs a two‑stage training process that first improves privilege following by jointly optimizing rollouts with and without privileged guidance, then selectively distills step‑specific reasoning and memory guidance via a privilege‑conditioned self‑teacher. Experiments on AndroidWorld and MobileWorld benchmarks demonstrate that GUI‑SD‑v2 outperforms existing OPSD baselines and state‑of‑the‑art methods in Pass@1 and Pass@3 success rates.
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
arXiv:2607. 04425v1 Announce Type: cross Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction.
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
arXiv:2608. 15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution.