arXiv AI

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

arXiv:2603. 18373v4 Announce Type: replace-cross Abstract: When VLMs answer correctly, do they genuinely rely on visual information?

arXiv Computer Vision
Sep 11

HALDETECT at ImageEval 2026 Shared Tasks: Answer-First Contrastive Grounding with QLoRA

HALDETECT is a system developed for the English hallucination-detection track of ImageEval 2026, where the task is to identify the single visually grounded statement among three culturally plausible options. The approach treats the problem as a contrastive decision, outputs the answer before an explanation, and bases reasoning on colour/texture, shape/form, and context. The best model fine‑tunes Qwen2.5‑VL‑7B‑Instruct with 4‑bit QLoRA, freezes the vision encoder, and achieves a Contrastive Instability score of 0.035 on the test set, placing third among eight teams.

By Syed Mohaiminul Hoque, Md Sakhawat Hossain
arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
Hugging Face Trending Papers
Aug 19

When Safety Overrides Vision: Exploring Dynamics between Vision Influence and Safety Alignment in Vision-Language Models

Aligned vision‑language models (VLMs) are designed to combine grounded visual reasoning with safe generation. The study finds that when safety constraints are applied, these models often abstain from answering questions that they could answer under default instruction, yet visual evidence continues to influence the decoding process. The authors show that safety‑induced abstention alters late‑stage hidden‑state dynamics, and that targeted interventions can restore grounded answering without retraining or changing visual inputs.