← Back to all news
arXiv Computer Vision August 24, 2026 By Junqi Wu, Kaihua Tang, Xuanwen Chen, Hongzhi Li, Jianqiang Huang, Xian-Sheng Hua

AffordAny: Open-World 3D Affordance Grounding from Monocular RGB Images via Vision-Language-Guided Geometric Reasoning

Read the original on arXiv Computer Vision →

The Flow has not summarised this story yet — read it at arXiv Computer Vision.

  • llms
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 30

GROW$^2$: Grounding Which and Where for Robot Tool Use

arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?

By Yuhong Deng, Yuyao Liu, David Hsu
llmsagentsroboticsmultimodalbenchmarks
More like this →
arXiv AI
Jul 8

OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding

arXiv:2512. 23020v3 Announce Type: replace-cross Abstract: 3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes.

By Wenyuan Huang, Zhenyu Zhang, Zhao Wang, Zhou Wei, Ting Huang, Fang Zhao, Jian Yang
llmsbenchmarks
More like this →
arXiv AI
Jul 7

Token-Based Affordance Grounding with Large Vision-Language Models

arXiv:2607. 03595v1 Announce Type: cross Abstract: Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception.

By Seung Il Lee, Qinqian Lei, Daguang Xu, Dong Yang, Robby T. Tan, Yixin Chen, Bo Wang
llmsroboticsmultimodalbenchmarks
More like this →
arXiv AI
Jul 8

Harnessing Generative Image Models for Training-Free Primitive Shape Abstraction

arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.

By Gregor Kobsik, Tim Elsner, Leif Kobbelt
llmscomputer-visionroboticsfine-tuningmultimodal
More like this →
arXiv AI
Jul 1

PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding

arXiv:2606. 31148v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions.

By Duc Cao Dinh, Khai Le-Duc, Florent Draye, Chris Ngo, Terry Jingchen Zhang, Bernhard Sch\"olkopf, Zhijing Jin
llmsefficiencymultimodalbenchmarks
More like this →
arXiv Computer Vision
5d ago

Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding

arXiv:2608.21136v1 Announce Type: new Abstract: Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supe...

By Jie Xu, Na Zhao
llmsagentscomputer-visionrobotics
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea