arXiv Computer Vision By Wonbin Son, Gyumun Choi, Junil Seo, Seungmin Rho, Mi Young Lee, Hyungjoon Kim

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

Read the original on arXiv Computer Vision →

The paper investigates when frozen vision‑language models (VLMs) respond to edits made at the representation level, specifically by modifying object‑token sets instead of the raw image. It introduces an answer‑key‑free protocol that evaluates edits without annotated post‑edit answers, revealing that VLM responses depend on explicit edit teaching, token cleanliness, density, and a separable reading axis. The study demonstrates that image‑free token edits can preserve most of the performance of matched patch‑token baselines and even outperform oracle methods on the VRSBench dataset across multiple remote‑sensing datasets and model backbones.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Sep 3

When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Reveals

The paper investigates how frozen vision‑language models (VLMs) respond to edits made directly to their internal object‑token representations, bypassing the image input. It introduces an answer‑key‑free protocol that evaluates edits by logical consistency and self‑audit, revealing that responses depend on explicit edit teaching, token cleanliness, density, and a separable reading axis. The study demonstrates that image‑free token edits can preserve most free‑text VQA performance and even outperform oracle methods on remote‑sensing benchmarks, with findings consistent across multiple datasets and model backbones.

arXiv AI
Jul 15

Visual Access Boundaries in Vision-Language Model Reasoning

arXiv:2607. 12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces.

By Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo
Hugging Face Trending Papers
Jul 14

Visual Access Boundaries in Vision-Language Model Reasoning

Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extended when VLMs generate longer reasoning traces. We ask whether CoT requires continued access to image tokens, or whether it mainly operates over visual information already made available earlier in the forward pass.

arXiv AI
Sep 17

AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon evaluation framework for vision‑language models that tests hidden‑target identification across galleries of visually similar images. The benchmark requires a model to ask a sequence of attribute‑focused Yes/No questions, receiving Yes/No/Unsure feedback and penalizing invalid actions with Skip, thereby stressing question selection under uncertainty, constraint tracking, and fine‑grained discrimination. Using the Guess My Preferred Dress task, the study shows that final‑answer accuracy alone overstates performance, as models may guess correctly without verified evidence, waste turns, or violate the protocol, highlighting the need for combined visual discrimination, informative questioning, and robust protocol adherence.

By Min Wang, Ata Mahjoubfar
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li