arXiv AI
Sep 2

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

The paper investigates a failure mode in multimodal large language models called multimodal contextual sycophancy, where external text can override conflicting image evidence. A diagnostic set of 998 cases independently varies visual evidence, commonsense priors, and external text to probe when this failure occurs. Experiments across six models show that a System‑2 Visual Arbitration (S2VA) approach, which withholds text from the visual witness, significantly improves performance over direct witness reports, with the best information boundary varying by model and context source.

By Yi-Cheng Lai, Hen-Hsen Huang