arXiv:2609.06011v1 Announce Type: cross
Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexp...
By Yen-Ting Piao, Shu-Yun Chen, Chin-Hui Chu, Chun-Wei Chen, Shih-Yun Shan Kuan, Hung-yi Lee, Yun-Nung Chen
arXiv:2609.00293v1 Announce Type: new
Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in contex...
By Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick
The paper investigates a failure mode in multimodal large language models called multimodal contextual sycophancy, where external text can override conflicting image evidence. A diagnostic set of 998 cases independently varies visual evidence, commonsense priors, and external text to probe when this failure occurs. Experiments across six models show that a System‑2 Visual Arbitration (S2VA) approach, which withholds text from the visual witness, significantly improves performance over direct witness reports, with the best information boundary varying by model and context source.
By Yi-Cheng Lai, Hen-Hsen Huang
arXiv:2608. 04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer.
By De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
By Nan Li, Albert Gatt, Massimo Poesio
arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.
By Danae S\'anchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott
arXiv:2606. 19965v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts.
By Yihao Wang, Zijian He, Jie Ren, Keze Wang
The paper introduces visual adaptations of counterfactual tests—vCT and vCCT—to evaluate whether chain-of-thought explanations in vision‑language models faithfully reflect the visual evidence driving predictions. Using these tests, the authors benchmark eight open‑source VLMs on two datasets and find that CoTs often fail to track visual evidence, sometimes omitting removed objects or mentioning them inconsistently. They also release two new datasets, Counter‑SNLI‑VE and Counter‑A‑OKVQA, consisting of image pairs that differ by a single object to facilitate further research.
By Bayar Menzat, Maximilian S\"uss, Ruizhi Wang, Benno Steinegger, Thomas Lukasiewicz, Oana-Maria Camburu
The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.
By Jungyeon Lee, Yejin Yoon, Taeuk Kim
arXiv:2509. 22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect.
By Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.
By Hong-Han Wang, Yuntao Wang, Hu Ding
arXiv:2604.04692v3 Announce Type: replace-cross
Abstract: Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-o...
By Jaeyoon Jung, Yejun Yoon, Kunwoo Park