GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
arXiv:2608. 03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents.
arXiv:2511. 00810v4 Announce Type: replace-cross Abstract: Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable regions on the screen.
arXiv:2608. 03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents.
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input images into static feature representations and perform reasoning primarily within the linguistic domain, often overlooking the rich perceptual cues and explicit spatial geometry inherent in images.
arXiv:2608. 09654v1 Announce Type: new Abstract: GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots.
GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates.
arXiv:2506. 01850v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs).
arXiv:2607. 10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks.
arXiv:2607. 27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning.
arXiv:2606. 16122v1 Announce Type: new Abstract: Visual thinking should not only sound right; it should show its evidence.
arXiv:2608. 11191v1 Announce Type: cross Abstract: GUI Visual Grounding is a fundamental capability for GUI agents.
arXiv:2608. 02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current models struggle with accurate visual grounding and coherent reasoning chains.
arXiv:2603. 00171v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) are shifting towards "Thinking with Images" by actively exploring image details.
arXiv:2511. 07332v2 Announce Type: replace-cross Abstract: Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements.