RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608. 03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents.
IVSGround introduces a lightweight view selector that learns to choose the most informative camera views for vision‑language model (VLM) based 3D visual grounding, replacing heuristic view selection. The selector is trained via a two‑stage rejection sampling process that uses feedback from a reasoning VLM to generate supervision signals. Experiments on ScanRefer and NR3D demonstrate that IVSGround consistently improves grounding accuracy over existing zero‑shot pipelines, underscoring the importance of selecting where to look for effective 3D visual grounding.
arXiv:2511. 00810v4 Announce Type: replace-cross Abstract: Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable regions on the screen.
arXiv:2609.19143v1 Announce Type: cross Abstract: Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language mode...
arXiv:2608.22429v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...
arXiv:2608.30233v1 Announce Type: new Abstract: Generalized Visual Grounding (GVG) task aims to localize targets in an image based on referring expressions, extends the classical visual grounding par...