arXiv Computer Vision

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

The paper investigates when frozen vision‑language models (VLMs) respond to edits made at the representation level, specifically by modifying object‑token sets instead of the raw image. It introduces an answer‑key‑free protocol that evaluates edits without annotated post‑edit answers, revealing that VLM responses depend on explicit edit teaching, token cleanliness, density, and a separable reading axis. The study demonstrates that image‑free token edits can preserve most of the performance of matched patch‑token baselines and even outperform oracle methods on the VRSBench dataset across multiple remote‑sensing datasets and model backbones.

Hugging Face Trending Papers
Sep 3

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

The paper investigates how frozen vision‑language models (VLMs) respond to edits made directly to their internal object‑token representations, bypassing the image input. It introduces an answer‑key‑free protocol that evaluates edits by logical consistency and self‑audit, revealing that responses depend on explicit edit teaching, token cleanliness, density, and a separable reading axis. The study demonstrates that image‑free token edits can preserve most free‑text VQA performance and even outperform oracle methods on remote‑sensing benchmarks, with findings consistent across multiple datasets and model backbones.

arXiv AI
Jul 15

Visual Access Boundaries in Vision-Language Model Reasoning

arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.

By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.

arXiv AI
Sep 17

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.

By Min Wang, Ata Mahjoubfar
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv Computer Vision
Sep 14

Semantically Aligned Gradient-Driven Context-Preserving Image Editing

Semantically Aligned Gradient-Driven Context-Preserving Image Editing (IABEdit) is a model‑agnostic framework that embeds differentiable semantic verification into the training of generative image editors. By using a frozen vision‑language model to extract spatially‑aware descriptors from ground‑truth edits and a trainable aligner to reproduce them from generated outputs, the residual becomes a gradient that teaches the generator both what to edit and where, without adding inference‑time VLM cost. IABEdit is compatible with various backbones (e.g., U‑Net in Stable Diffusion and MMDiT in FLUX) and improves structural fidelity on MagicBrush, achieves state‑of‑the‑art instruction adherence on RealEdit and EMU Edit, and outperforms the proprietary Gemini agent on the D‑LORD surveillance benchmark under heavy occlusion. "whyItMatters":"IABEdit demonstrates that incorporating semantic verification during training can produce more accurate, well‑localized edits and outperform existing methods even in challenging surveillance scenarios, as shown by its superior metrics and human/GPT‑4o evaluations."

By Chiranjeev Chiranjeev, Muskan Dosi, Mayank Vatsa, Richa Singh
arXiv Computer Vision
Sep 14

Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap

ReVisIT is a train‑free framework that turns retrieved image‑label pairs into units of visual thought, combining structured class definitions, multimodal retrieval, and alternating user/assistant injection before joint decoding. On several benchmarks—including Fast Open MiniImageNet, Bongard‑OpenWorld, and the newly released MAAC‑Bench—ReVisIT achieves performance comparable to or surpassing large, trained models while using far fewer parameters. The approach demonstrates that high‑quality retrieval and a simple turns layer can provide a universal performance boost across diverse multimodal tasks.

By Bingchen Huang, Zhiling Wang, Yifu Chen, Yuanchao Du
arXiv Computation and Language
Sep 25

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

The paper introduces a perception interface that separates vision from language in vision‑language models. A frozen perception stack detects objects, a deterministic semantic serializer converts the perceived state into text, and a standard text‑only large language model (LLM) answers questions. Experiments on a campus‑robot benchmark show that this serialized interface outperforms a zero‑shot VLM of the same language‑model size, especially as the language model shrinks, and that the advantage persists under paraphrase and different supervision regimes.

By Cong Xu, Ravi Sankar