arXiv AI By Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang

Token-Based Affordance Grounding with Large Vision-Language Models

Read the original on arXiv AI →

arXiv:2607. 03595v1 Announce Type: cross Abstract: Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 23

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Traditional methods typically encode input images into static feature representations and perform reasoning primarily within the linguistic domain, often overlooking the rich perceptual cues and explicit spatial geometry inherent in images.