arXiv:2610.01243v1 Announce Type: cross
Abstract: Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such...
By Huichan Seo
arXiv:2610.00111v1 Announce Type: new
Abstract: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal rea...
By Rasul Khanbayov, Hasan Kurban
arXiv:2608. 13267v1 Announce Type: cross Abstract: Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading).
By Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee, Christian Greisinger, Steffen Eger, Pius Onobhayedo, Wei Zhao
The paper demonstrates that vision‑language models’ refusal behavior is heavily influenced by whether an image is attached to a request, even when the image is blank or unreadable. Attaching such an image shifts refusal scores by large margins for borderline‑benign prompts while leaving genuinely neutral instructions largely unchanged. This effect varies with image properties, persists across checkpoints, and is not mitigated by explicit instructions to ignore the image.
By Haoyu Zhang, Yi Feng, Hanwen Liu, Shibo Zheng, Zhuoxi Wang, Yang Chen, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2609.00868v1 Announce Type: cross
Abstract: Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual inp...
By Genpei Zhang
The paper shows that the wording of prompts in vision‑language models (VLMs) can either improve or worsen robustness to image corruption. Verbose prompts broaden the cross‑modal attention’s frequency filter, making the model less sensitive to corruptions, while semantically complex prompts narrow the filter and increase vulnerability. Experiments on Qwen3‑VL and LLaVA‑OneVision confirm that adding padding or verbose phrasing reduces answer drift by 70–81% on 8B models.
By Farooq Ahmad Wani, Maria Sofia Bucarelli, Mujtaba Hussain Mirza, Oleksandr Pryymak, Aryo Pradipta Gema, Iacopo Masi, Pasquale Minervini, Fabrizio Silvestri
arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.
By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li
arXiv:2609.26093v1 Announce Type: new
Abstract: Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-g...
By Feixiang Liu, Qiang Qiu, Qingyang Li, Hui Xu
VisionQ is a new benchmark for qualitative analysis in computer vision that evaluates vision‑language models (VLMs) on criterion‑conditioned visual discrimination. It is built from over 1,800 peer‑reviewed comparison figures in CVPR and ICCV papers, linking each image crop to author‑stated visual claims through 3,911 hand‑annotated data points. The benchmark includes a 51‑leaf taxonomy of visual criteria, a protocol that hides method identities and reports accuracy per criterion, and a DPO‑tuned Gemma‑4‑E4B judge that improves accuracy on a held‑out test set.
By Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, Hoang D. Nguyen
arXiv:2609.15635v1 Announce Type: cross
Abstract: A radiology report can already answer a clinical question, so it is hard to tell whether a vision-language model also uses the image. ModaLens, a pai...
By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Maximin Lange, Quang Bui, Anqi Peter Li, Felipe Ocampo Osorio, Rafi Al Attrach, Kushul Reddy Palakala, Sahil Kapadia, Zakaria Laouabdia Sellami, Xinyue Zhang, Ashley Zhang, Leo Anthony Celi
Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.
By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
arXiv:2607. 15565v1 Announce Type: cross Abstract: Where should the question go in a vision-language model (VLM) prompt: before the image or after it?
By Rakshanda Hassan Abhinandan, John Galeotti, Deva Ramanan, Gautam Rajendrakumar Gare