arXiv AI By Xinyu Qiu, Yunzhu Zhang, Heng Jia, Shuheng Shen, Changhua Meng, Linchao Zhu

VISTA: View-Consistent Self-Verified Training for GUI Grounding

Read the original on arXiv AI →

arXiv:2606. 14579v1 Announce Type: new Abstract: When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult instances or all successes on easy ones, yielding no useful relative advantage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 27

PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

PointRL introduces a verifiable reinforcement learning framework that learns point-level vision‑language grounding from heterogeneous annotation evidence such as bounding boxes, masks, and instance labels. The method converts these annotations into pointing instructions while preserving target supports, instance membership, and set constraints as hidden verifier evidence, which a deterministic checker uses to score predictions. Evaluation on PointArena shows that PointRL improves Qwen3.5‑4B’s accuracy from 56.11% to 65.58%, and similar gains are observed on RoboSpatial, BLINK, and Ref‑Adv benchmarks.

By Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang
arXiv AI
Jul 2

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

arXiv:2601. 14232v2 Announce Type: replace-cross Abstract: Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systematic analysis.

By Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov