arXiv AI By Vanshika Vats, Ashwani Rathee, James Davis

InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation

Read the original on arXiv AI →

InsightSeg introduces an episodic memory system that captures successful correction episodes from multi-agent refinement processes and transforms them into reusable, visually grounded insights. These insights are distilled into natural-language directives and linked to specific image patches via visual concept vectors, allowing the segmentation agent to anticipate and avoid guideline-specific errors on new images. The approach improves both first-pass and final segmentation quality on Waymo and Cityscapes datasets while reducing the number of required refinement steps.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 21

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

arXiv:2607. 16311v1 Announce Type: cross Abstract: Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.

By Jingyu Sun, Jiachen Tu, Yuyang Xue, Yaoxin Jiang, Guoyi Xu, Zhengtao Yao, Rui Qian, Yizheng Sun, Hongpeng Zhou, Jingyuan Sun, Yan Lin
arXiv AI
3d ago

Rethinking Multi-Image Re-Representation in Multi-Image Understanding

The paper introduces Mosaic, a multi-image visual harness that lets large language‑vision models (MLLMs) construct visual intermediates using ten composable image operations. It evaluates five re‑representation settings on existing multi‑image benchmarks and a new grounding‑focused benchmark, MosaicBench, finding that visual re‑representation benefits tasks requiring precise visual evidence more than those dominated by high‑level semantics. MosaicAgent‑8B is trained via reinforcement learning to compose these operations without demonstration trajectories, demonstrating diverse problem‑solving patterns.

By Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp
arXiv AI
Sep 4

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.

By Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu
arXiv Computer Vision
Sep 22

0.5\%>100\%: Bidirectional Reciprocal Learning for Referring Image Segmentation

The paper introduces Bidirectional Reciprocal Learning (BRL), a parameter‑efficient fine‑tuning framework for referring image segmentation that operates on frozen vision foundation models. BRL employs two lightweight adapters—Reciprocal Attention Adapter (RAA) for token‑level cross‑modal attention and Reciprocal Gate Adapter (RGA) for channel‑level gating—to enable hierarchical, bidirectional information flow between vision and language. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show that BRL outperforms existing methods while updating fewer than 0.5% of backbone parameters.

By Xiaoqiang Lu, Licheng Jiao, Lingling Li, Yuting Yang, Long Sun, Wenping Ma, Xu Liu, Fang Liu