arXiv AI

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

The paper investigates how different multimodal design choices affect the performance of misinformation detection systems. Using over 3,375 experiments across three benchmark datasets and various pre‑trained vision and language models, the authors systematically compare design options and conduct robustness analyses. The study offers practical guidance on which choices improve detection, when they may fail silently, and which pipeline components most influence model behavior, addressing four key research questions.

arXiv Computation and Language
Sep 2

Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict

The paper investigates how multimodal large language models (MLLMs) handle conflicting evidence presented in text, image, or both forms. Across 13 MLLMs and two datasets, the authors find that models are not robust to knowledge conflict: they tend to accept contradictory image evidence more readily than contradictory text, and when both modalities conflict the preference is arbitrary, depending on input order, model, and dataset. The instability degrades multimodal retrieval-augmented generation and can be exploited by adversarial attacks, while simple mitigation techniques such as prompting, steering, and direct preference optimization largely fail, with supervised fine‑tuning offering only moderate improvement.

By Jungyeon Lee, Yejin Yoon, Taeuk Kim
arXiv AI
4d ago

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

The paper investigates whether multimodal large language models (MLLMs) can generate and detect realistic multimodal fake news on social media. Using a multi‑agent framework—comprising a story agent, an image agent, and a critic agent—the authors produced over 9,000 paired multimodal news posts across science, health, and entertainment domains. They benchmarked 16 open‑ and closed‑source MLLMs for automated detection and found that most models fall far short of human accuracy, especially in identifying image authenticity, highlighting the need for stronger defenses against social media fake news.

By Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, Xiuzhen Zhang