arXiv:2609.16646v1 Announce Type: new
Abstract: When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instrument...
By Zhipeng Zhao, Wenxu Wang, Peishun Liu, Ruichun Tang
arXiv:2609.26093v1 Announce Type: new
Abstract: Vision-language models can answer spatial relation questions confidently even when the image supports an incompatible relation. We formulate relation-g...
By Feixiang Liu, Qiang Qiu, Qingyang Li, Hui Xu
arXiv:2608. 19807v1 Announce Type: new Abstract: Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall model performance against annotated ground truth.
By Rongyu Yu, Ke Niu, Fengxiang He
arXiv:2606. 06890v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence.
By Runyu Zhou, Qi Zhang, Qixun Wang, Yisen Wang
arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.
By Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu
HALDETECT is a system developed for the English hallucination-detection track of ImageEval 2026, where the task is to identify the single visually grounded statement among three culturally plausible options. The approach treats the problem as a contrastive decision, outputs the answer before an explanation, and bases reasoning on colour/texture, shape/form, and context. The best model fine‑tunes Qwen2.5‑VL‑7B‑Instruct with 4‑bit QLoRA, freezes the vision encoder, and achieves a Contrastive Instability score of 0.035 on the test set, placing third among eight teams.
By Syed Mohaiminul Hoque, Md Sakhawat Hossain
arXiv:2609.38362v1 Announce Type: new
Abstract: Generative vision-language models (VLMs) such as Qwen-VL and LLaVA achieve strong zero-shot performance on tasks overlapping with their pretraining dis...
By Hung-Jen Chen, Yu-Heng Ho, Ting-Yao Huang, Po-Hsiang Hsu, Li-Yu Chen, Chun-Yi Lee, Min Sun
arXiv:2608. 13167v1 Announce Type: cross Abstract: When visual evidence is occluded or chaotic, models should abstain.
By Fnu Pramono, John Cai, Sourabh Kulkarni
arXiv:2609.09184v1 Announce Type: new
Abstract: Vision-language model (VLM) confidence may change in aggregate when visual evidence is degraded while remaining structurally inconsistent within indivi...
By Muhamathu Ameer Ali Aacaas Muhamath
arXiv:2602.06652v2 Announce Type: replace
Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...
By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
Aligned vision‑language models (VLMs) are designed to combine grounded visual reasoning with safe generation. The study finds that when safety constraints are applied, these models often abstain from answering questions that they could answer under default instruction, yet visual evidence continues to influence the decoding process. The authors show that safety‑induced abstention alters late‑stage hidden‑state dynamics, and that targeted interventions can restore grounded answering without retraining or changing visual inputs.
Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these latents and their visual targets, i.