It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2610.01243v1 Announce Type: cross Abstract: Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such...
arXiv:2610.00111v1 Announce Type: new Abstract: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal rea...
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
The paper demonstrates that vision‑language models’ refusal behavior is heavily influenced by whether an image is attached to a request, even when the image is blank or unreadable. Attaching such an image shifts refusal scores by large margins for borderline‑benign prompts while leaving genuinely neutral instructions largely unchanged. This effect varies with image properties, persists across checkpoints, and is not mitigated by explicit instructions to ignore the image.
arXiv:2609.00868v1 Announce Type: cross Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.