RankGround is a two‑stage framework for GUI grounding that uses a single Vision‑Language Model call per query. It introduces GroundRanker, a lightweight multimodal reranker that selects the most promising crop from a dense candidate set, trained with a two‑stage curriculum on ranking supervision data derived from existing grounding datasets. Experiments show RankGround outperforms strong baselines, achieving 1.4× faster inference and a 5.5% average improvement in localization accuracy over the second‑best method across all backbones and screen scales.
By Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
arXiv:2511. 00810v4 Announce Type: replace-cross Abstract: Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable regions on the screen.
By Shijie Zhou, Viet Dac Lai, Hao Tan, Jihyung Kil, Wanrong Zhu, Changyou Chen, Ruiyi Zhang
arXiv:2608. 09654v1 Announce Type: new Abstract: GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots.
By Yuke Li, Xuehan Hou
GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates.
arXiv:2604. 14262v2 Announce Type: replace-cross Abstract: GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming.
By Yangyue Wang, Harshvardhan Sikka, Yash Mathur, Tony Zhou, Jinu Nyachhyon, Pranav Guruprasad
arXiv:2608. 11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents.
By Shiyu Xuan, Zechao Li