The paper introduces OpenRef, a benchmark for Referring Expression Comprehension (REC) designed for open‑world scenarios. OpenRef expands beyond simple settings by including diverse visual domains, variable target counts (multi‑target and none‑target), and a rich vocabulary with proper nouns, polysemous words, and ordinal terms. It also proposes new evaluation metrics—F1 for grounding accuracy and N3R for negative expression rejection—and presents a training‑free Multi‑task Consistency Checker (MCC) that improves model performance with a single click.
By Zongjian Wu, Lei Zhang
arXiv:2609.24510v2 Announce Type: replace
Abstract: Recent advances in vision foundation models (VFMs) have shown remarkable capabilities across diverse unimodal visual tasks. However, adapting VFMs...
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.
By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu
The paper introduces Lens, a training‑free framework that aligns multimodal representations with the semantic perspective required by downstream tasks. Lens uses a task‑specific readout phrase to anchor the perspective and then aggregates token states after the full input, ensuring the extracted representation reflects task‑conditioned evidence integration rather than generic salient content. The method achieves a Precision@1 of 63.9 across 36 MMEB datasets, outperforming the nearest training‑free baseline by 10.2 points.
By Xinran Liu, Shouqian Shi, Yixian Chen, Ruizhi Chen, Xin-Wei Yao, Sheng Zhong
arXiv:2608. 04698v1 Announce Type: cross Abstract: We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it does not (negative sample).
By Xuzheng Yang, Jun Ling, Tao Huang, Caiyan Qin, Peng Wang
Referring Expression Segmentation (RES) aims to generate a pixel-level mask for the object specified by a language expression. Recent methods based on multimodal large language models (MLLMs) often re...
arXiv:2505.22850v3 Announce Type: replace
Abstract: Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle t...
By Kostas Triaridis, Panagiotis Kaliosis, E-Ro Nguyen, Jingyi Xu, Dimitris Samaras, Hieu Le
DRAgent introduces a discriminative reasoning framework for Referring Expression Segmentation that replaces direct coordinate prediction with a two‑stage selection process. First, a detector generates candidate boxes; then a multimodal large language model (MLLM) discriminates the correct target through high‑recall screening and instance‑wise verification. The chosen box serves as a spatial prompt for a foundation segmentation model, and a self‑consistency‑filtered data pipeline is used for LoRA‑based fine‑tuning, yielding competitive results on RefCOCO, RefCOCO+, and RefCOCOg.
By Yujie Qi, Luyan Zhang
arXiv:2608.22429v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this relian...
By Changjiang Jiang, Qiannian Zhao, Lei Xin, Jinxiang Xie, Preslav Nakov, Zhuohan Xie
arXiv:2606. 03564v2 Announce Type: replace-cross Abstract: Reasoning segmentation aims to segment target objects described by complex language through joint visual-textual reasoning.
By Yifan Cao, Xiaocui Yang, Faxian Wan, Shi Feng, Daling Wang, Yifei Zhang
arXiv:2501.04001v4 Announce Type: replace
Abstract: This work presents Sa2VA, the first comprehensive, unified model for dense grounded understanding of both images and videos. Unlike existing multi-...
By Haobo Yuan, Xiangtai Li, Tao Zhang, Yueyi Sun, Zilong Huang, Shilin Xu, Shunping Ji, Yunhai Tong, Lu Qi, Jiashi Feng, Ming-Hsuan Yang
arXiv:2608. 03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction.
By Haoqian Kang, Liupeng Li, Kuofeng Gao, Jinpeng Wang, Zhenyu Lu, Bin Chen, Ke Chen, Yaowei Wang