arXiv:2604. 14262v2 Announce Type: replace-cross Abstract: GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial reasoning rather than direct element naming.
By Yangyue Wang, Harshvardhan Sikka, Yash Mathur, Tony Zhou, Jinu Nyachhyon, Pranav Guruprasad
arXiv:2608. 11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents.
By Shiyu Xuan, Zechao Li
arXiv:2608. 03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents.
By Zichuan Fu, Shirong Wang, Wenlin Zhang, Guojing Li, Yimin Deng, Jingtong Gao, Junjia Qi, Hanyu Yan, Yefeng Zheng, Xiaopeng Li, Wanyu Wang, Xian Wu, Xiangyu Zhao
PointRL introduces a verifiable reinforcement learning framework that learns point-level vision‑language grounding from heterogeneous annotation evidence such as bounding boxes, masks, and instance labels. The method converts these annotations into pointing instructions while preserving target supports, instance membership, and set constraints as hidden verifier evidence, which a deterministic checker uses to score predictions. Evaluation on PointArena shows that PointRL improves Qwen3.5‑4B’s accuracy from 56.11% to 65.58%, and similar gains are observed on RoboSpatial, BLINK, and Ref‑Adv benchmarks.
By Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang
arXiv:2606. 13156v2 Announce Type: replace-cross Abstract: Letting a vision-language model (VLM) think longer at test time has driven much recent progress.
By Animesh Tripathy, Aswanth Krishnan
arXiv:2601. 14232v2 Announce Type: replace-cross Abstract: Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systematic analysis.
By Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov