arXiv Machine Learning By Tianze Yang, Yucheng Shi, Ruitong Sun, Ninghao Liu, Jin Sun

Self-Improving Small Object Grounding in LVLMs

Read the original on arXiv Machine Learning →

arXiv:2606. 01612v1 Announce Type: cross Abstract: Can internal attention patterns in Large Vision Language Models (LVLMs) identify reliable small-object boxes without fine-tuning?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 7

Where to Look Matters: Learning Influential Views for VLM-based 3D Visual Grounding

IVSGround introduces a lightweight view selector that learns to choose the most informative camera views for vision‑language model (VLM) based 3D visual grounding, replacing heuristic view selection. The selector is trained via a two‑stage rejection sampling process that uses feedback from a reasoning VLM to generate supervision signals. Experiments on ScanRefer and NR3D demonstrate that IVSGround consistently improves grounding accuracy over existing zero‑shot pipelines, underscoring the importance of selecting where to look for effective 3D visual grounding.

By Tsung-Chih Chiang, Hsuan-Kung Yang, Jou-Min Liu, Ting-Ru Liu, Chun-Wei Huang, Quan Kong, Chun-Yi Lee