Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.06011v1 Announce Type: cross Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexp...
arXiv:2609.00293v1 Announce Type: new Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in contex...
The paper investigates a failure mode in multimodal large language models called multimodal contextual sycophancy, where external text can override conflicting image evidence. A diagnostic set of 998 cases independently varies visual evidence, commonsense priors, and external text to probe when this failure occurs. Experiments across six models show that a System‑2 Visual Arbitration (S2VA) approach, which withholds text from the visual witness, significantly improves performance over direct witness reports, with the best information boundary varying by model and context source.
arXiv:2608. 04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer.
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.