Vision Harnessing Agent for Open Ad-hoc Segmentation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.
VTOS (Vision Tools Orchestration Search) is a framework that adaptively orchestrates vision foundation tools—such as open‑vocabulary detectors, segmentation models, and post‑processing operators—by jointly searching for executable solution programs and observer programs that diagnose failures and provide feedback. The observer programs feed observations into a shared VisionThoughts knowledge base, guiding subsequent searches. In two case studies—dense object counting on LVIS‑Count and zero‑shot plant‑disease segmentation on PlantSeg‑OOD—VTOS outperforms static tool pipelines and agentic visual‑programming baselines, especially in complex scenarios like dense, occluded scenes and out‑of‑distribution segmentation.
arXiv:2609.24362v1 Announce Type: new Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and persistent files, yet their extension from language m...
arXiv:2608.21762v1 Announce Type: cross Abstract: Vision-language models (VLMs) fail many detail-centric questions for a concrete reason: the answer is visible in the image, yet lost after the image...
InsightSeg introduces an episodic memory system that captures successful correction episodes from multi-agent refinement processes and transforms them into reusable, visually grounded insights. These insights are distilled into natural-language directives and linked to specific image patches via visual concept vectors, allowing the segmentation agent to anticipate and avoid guideline-specific errors on new images. The approach improves both first-pass and final segmentation quality on Waymo and Cityscapes datasets while reducing the number of required refinement steps.
The paper introduces Visual Retrieval Heads (VRHs), a small fraction of attention heads in vision‑language models that are causally responsible for grounding text descriptions to image regions. By recasting head‑scoring methods and evaluating across eleven VLMs and five benchmarks, the authors show that masking the top 20 VRHs can drop grounding accuracy by up to 80 percentage points, while random masking has little effect. VRHs generalize across various visual reference tasks, preserve output format while corrupting localization, and transfer causally across models sharing an LLM backbone.