Learning Reliable GUI Agents under Imperfect Priors
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
ScreenSearch is a system for exploring desktop GUI states under partial observability. It combines structural screen retrieval, deduplication, and an ambiguity-aware PUCT graph-bandit to expand the reachable frontier while reducing state ambiguity. Across 11 applications, it collected over 1 M screenshots and 30 K deduplicated states, demonstrating that both ambiguity reduction and frontier exploration are essential for effective corpus building.
arXiv:2606. 29705v1 Announce Type: new Abstract: Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models.
arXiv:2606. 12817v1 Announce Type: new Abstract: Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension.
arXiv:2606. 12817v2 Announce Type: replace Abstract: Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension.
GUI-CC is a benchmark designed to assess the contextual consistency of GUI world models when used as agent environments, rather than just one‑step next‑screen predictors. It includes two tracks: an offline reference‑action track that rolls models along real mobile GUI trajectories, and an online agent‑loop track where fixed probing agents interact with model‑generated UIs. The benchmark evaluates transition fidelity, plausibility, contextual consistency, and task progress across 500 offline trajectory tasks and 200 online tasks spanning 30 mobile apps.
The paper introduces AnTrap, a benchmark that injects dynamic perturbations into Android GUI agent execution to evaluate robustness against runtime anomalies. It presents a taxonomy of anomalies across four layers—State, Thinking, Action, and Round—with ten subcategories, and a pipeline that maintains task solvability while adding realistic adversarial conditions. Experiments on 16 leading GUI models show universal vulnerability, and reinforcement learning can mitigate some traps but not deep contextual ones like state deadlock.