arXiv Computer Vision

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.

arXiv Computer Vision
Aug 25

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv:2501.04001v4 Announce Type: replace Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...

By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv Computer Vision
Sep 3

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

The paper presents CycleGRPO, a reinforcement learning framework that unifies region understanding and localization for multimodal large language models (MLLMs). By treating the MLLM as both actor and critic, the method generates region captions and immediately grounds them back into spatial coordinates, using a token‑level cycle‑consistency reward that obviates the need for textual ground truths. Experiments on SAMTok demonstrate that CycleGRPO can bootstrap region captioning, VQA, grounded dialogue, and referring segmentation simultaneously, achieving consistent performance gains without task‑specific fine‑tuning.

By Xin Zhang, Haochen Wang, Yikang Zhou, Zhuochen Wang, Xiangtai Li, Robby T. Tan
arXiv Computer Vision
Sep 10

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

arXiv:2609.09396v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...

By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
arXiv Computer Vision
Sep 18

Region-Level Policy Optimization for Fine-grained MLLM Perception

The paper introduces Vision‑RL2, a region‑level reinforcement learning approach that optimizes a lightweight proposal network for fine‑grained multimodal large language model (MLLM) perception. By treating coherent image regions as actions and scoring them with a frozen MLLM reader, the method selectively focuses visual resolution on evidence, reducing token usage while improving accuracy across multiple benchmarks and backbones. The approach eliminates the need for region annotations, response sampling, or reasoning trajectories, and the refined proposals enable sparse encoding that magnifies relevant evidence.

By Yuheng Shi, Xiaohuan Pei, Minjing Dong, Chang Xu
arXiv AI
Jun 30

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis

arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.

By Minh-Loi Nguyen, Nghiem Tuong Diep, Hung Khang Nguyen, Minh Le, Doanh Le Thien, Hoang H. Tran, Dung D. Le, Vu N. Duong, Daniel Sonntag, An Thai Le, Duy Minh Ho Nguyen, Vien Anh Ngo, Tran Van Nhiem