arXiv AI

Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

arXiv AI
Sep 2

Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy

The paper investigates a failure mode in multimodal large language models called multimodal contextual sycophancy, where external text can override conflicting image evidence. A diagnostic set of 998 cases independently varies visual evidence, commonsense priors, and external text to probe when this failure occurs. Experiments across six models show that a System‑2 Visual Arbitration (S2VA) approach, which withholds text from the visual witness, significantly improves performance over direct witness reports, with the best information boundary varying by model and context source.

By Yi-Cheng Lai, Hen-Hsen Huang
arXiv AI
Sep 10

Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

The paper introduces visual adaptations of counterfactual tests—vCT and vCCT—to evaluate whether chain-of-thought explanations in vision‑language models faithfully reflect the visual evidence driving predictions. Using these tests, the authors benchmark eight open‑source VLMs on two datasets and find that CoTs often fail to track visual evidence, sometimes omitting removed objects or mentioning them inconsistently. They also release two new datasets, Counter‑SNLI‑VE and Counter‑A‑OKVQA, consisting of image pairs that differ by a single object to facilitate further research.

By Bayar Menzat, Maximilian S\"uss, Ruizhi Wang, Benno Steinegger, Thomas Lukasiewicz, Oana-Maria Camburu
arXiv Computation and Language
Sep 2

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.

By Jungyeon Lee, Yejin Yoon, Taeuk Kim
arXiv Machine Learning
Sep 25

The Alignment Illusion in Multimodal Large Language Models

The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.

By Hong-Han Wang, Yuntao Wang, Hu Ding