arXiv AI

VISTA: View-Consistent Self-Verified Training for GUI Grounding

arXiv:2606. 14579v1 Announce Type: new Abstract: When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often become either all failures on difficult instances or all successes on easy ones, yielding no useful relative advantage.

arXiv Computer Vision
Aug 27

PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence

PointRL introduces a verifiable reinforcement learning framework that learns point-level vision‑language grounding from heterogeneous annotation evidence such as bounding boxes, masks, and instance labels. The method converts these annotations into pointing instructions while preserving target supports, instance membership, and set constraints as hidden verifier evidence, which a deterministic checker uses to score predictions. Evaluation on PointArena shows that PointRL improves Qwen3.5‑4B’s accuracy from 56.11% to 65.58%, and similar gains are observed on RoboSpatial, BLINK, and Ref‑Adv benchmarks.

By Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang
arXiv AI
Jul 2

KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning

arXiv:2601. 14232v2 Announce Type: replace-cross Abstract: Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are unchanged, but existing benchmarks entangle multiple sources of shift and hinder systematic analysis.

By Egor Cherepanov, Daniil Zelezetsky, Alexey K. Kovalev, Aleksandr I. Panov
arXiv Computation and Language
Sep 17

RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

RankGround is a two‑stage framework for GUI grounding that uses a single Vision‑Language Model call per query. It introduces GroundRanker, a lightweight multimodal reranker that selects the most promising crop from a dense candidate set, trained with a two‑stage curriculum on ranking supervision data derived from existing grounding datasets. Experiments show RankGround outperforms strong baselines, achieving 1.4× faster inference and a 5.5% average improvement in localization accuracy over the second‑best method across all backbones and screen scales.

By Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang
arXiv Computer Vision
Sep 11

Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

The paper introduces EgoGenEval, a new benchmark that assesses the physical consistency of visual generators under ego‑motion by measuring Camera Motion Grounding and Scene State Preservation across 1,400 cases and 2,360 target views. Experiments on 16 pose‑free generators and two pose‑conditioned references show that current models struggle to maintain both camera motion and scene state simultaneously. A follow‑up study using EgoGen‑Train demonstrates that pairwise supervision does not effectively improve both metrics together, suggesting the need for a trajectory‑centric training paradigm.

By Yilin Long, Chenming Zhu, Zitang Gou, Jingli Lin, Tai Wang