arXiv AI

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

arXiv:2608. 00410v2 Announce Type: replace Abstract: Human language is highly polysemous.

arXiv Computation and Language
Sep 2

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.

By Jungyeon Lee, Yejin Yoon, Taeuk Kim
arXiv AI
Sep 15

(How) Do MLLMs Report Bistable Images Like Humans?

The study investigates whether multimodal large language models (MLLMs) report bistable images, like the duck‑rabbit, in a manner similar to humans. Using the LLaVA family, researchers examined two dimensions: modulability (the influence of visual cues and linguistic priors) and exclusivity (whether responses commit to a single interpretation). Results show that both visual and linguistic manipulations shift reports in human‑consistent ways while maintaining predominantly exclusive responses, driven by competing image‑token representations and distinct bottom‑up and top‑down pathways.

By Ryota Takatsuki, Tomoki Doi, Amane Watahiki, Anil K. Seth, Hitomi Yanaka
arXiv Computer Vision
Aug 27

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

The paper investigates why multimodal large language models (MLLMs) struggle with vision‑centric tasks when visual evidence conflicts with pretrained language knowledge. Using image reconstruction and a new WhatIfVis benchmark, the authors show that MLLMs preserve coarse‑grained visual attributes but fail to consistently use them, and that supervised fine‑tuning and activation patching can improve controllability of visual context sensitivity. The study demonstrates that the main bottleneck lies in the models’ inability to reliably regulate their reliance on visual evidence rather than in visual perception itself.

By Jiaang Li, Chengzu Li, Zhaochong An, Yifei Yuan, Xi Liu, Serge Belongie, V\'esteinn Sn{\ae}bjarnarson