arXiv Computer Vision

Vision Harnessing Agent for Open Ad-hoc Segmentation

arXiv AI
Sep 4

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.

By Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu
arXiv Computation and Language
Sep 3

VTOS: Learning to Orchestrate Vision Tools by Co-Searching Solutions and Observers

VTOS (Vision Tools Orchestration Search) is a framework that adaptively orchestrates vision foundation tools—such as open‑vocabulary detectors, segmentation models, and post‑processing operators—by jointly searching for executable solution programs and observer programs that diagnose failures and provide feedback. The observer programs feed observations into a shared VisionThoughts knowledge base, guiding subsequent searches. In two case studies—dense object counting on LVIS‑Count and zero‑shot plant‑disease segmentation on PlantSeg‑OOD—VTOS outperforms static tool pipelines and agentic visual‑programming baselines, especially in complex scenarios like dense, occluded scenes and out‑of‑distribution segmentation.

By Jinchao Ge, Lingqiao Liu, Shuwen Zhao, Lei Wang
arXiv AI
Sep 3

InsightSeg: Reusing Correction Insights for Guideline-Consistent Segmentation

InsightSeg introduces an episodic memory system that captures successful correction episodes from multi-agent refinement processes and transforms them into reusable, visually grounded insights. These insights are distilled into natural-language directives and linked to specific image patches via visual concept vectors, allowing the segmentation agent to anticipate and avoid guideline-specific errors on new images. The approach improves both first-pass and final segmentation quality on Waymo and Cityscapes datasets while reducing the number of required refinement steps.

By Vanshika Vats, Ashwani Rathee, James Davis
arXiv Computer Vision
Aug 28

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.

By Chanho Park, Daehyeon Choi, Jihyun Lee, Minhyuk Sung
arXiv AI
Sep 4

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

ENEAS is a unified, text‑promptable method that simultaneously provides precise instance tracking and high‑quality segmentation, and enables open‑concept discovery of any instance named by a text query. It extends the SeC architecture with a text‑prompting adapter and temporal memory to maintain targets through disappearance and avoid drifting, while a semantic verification layer combines visual embedding matching with conditional VLM refinement to filter ontological errors. Designed for 3D reconstruction, ENEAS delivers robust semantic tracking and segmentation across videos, libraries, and unordered collections, distinguishing true instances from look‑alike doppelgangers.

By Javier del Pino (SperidLabs), Salvador Rodr\'iguez (SperidLabs), Alejandro Garabito (SperidLabs), Javier \'Alvarez (SperidLabs), Chema Garabito (SperidLabs)
arXiv AI
Aug 26

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

VideoHarness‑RSI explores how improving the executable context‑construction program alone can enhance long‑video understanding with frozen vision‑language models. By recursively searching for better harnesses—programs that select and structure video segments—using an outer‑loop proposer that learns from prior programs and execution traces, the method consistently outperforms weaker hand‑crafted baselines and further improves upon stronger ones. The resulting harnesses transfer to other long‑video benchmarks without additional search, demonstrating that executable context construction is a distinct, reusable optimization layer.

By Guoyang Xu, Hao Chen
arXiv AI
Jul 21

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs

arXiv:2607. 16311v1 Announce Type: cross Abstract: Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.

By Jingyu Sun, Jiachen Tu, Yuyang Xue, Yaoxin Jiang, Guoyi Xu, Zhengtao Yao, Rui Qian, Yizheng Sun, Hongpeng Zhou, Jingyuan Sun, Yan Lin
arXiv Machine Learning
Jun 3

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.

By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang
arXiv AI
Sep 1

LOCI: A Locator-Critic with Refinement Loop

LOC I (Locator‑Critic) is a training‑free framework that separates visual search from evidence verification in Vision‑Language Models. It uses a Locator agent to propose candidate visual evidence and a Critic agent to assess its relevance, engaging in an iterative refinement loop that progressively improves the evidence until it is sufficient to answer a question. The approach yields state‑of‑the‑art results on several complex visual benchmarks, boosting accuracy for both open‑weight models like Qwen3‑VL and proprietary models such as Gemini 2.5 Pro.

By Walid Bousselham, Mathilde Caron, Arsha Nagrani, Cordelia Schmid