arXiv AI By Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

Read the original on arXiv AI →

arXiv:2608. 00410v2 Announce Type: replace Abstract: Human language is highly polysemous.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 2

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.

By Jungyeon Lee, Yejin Yoon, Taeuk Kim