arXiv AI

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

arXiv:2606. 24112v1 Announce Type: new Abstract: Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors.

arXiv Computer Vision
6d ago

MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

MM-VeriAgent is a reinforcement‑learning framework that learns to verify multimodal misinformation by leveraging a specialized toolkit called MM-VeriTools. The toolkit encapsulates the strongest models for textual, visual, and cross‑modal forgery analysis as callable tools with a unified interface. To improve training efficiency, the authors introduce a Tool‑Execution Cache that pre‑executes candidate tool calls and reuses cached outputs, resulting in substantial accuracy gains on MMFakeBench and reduced online tool executions during training.

By Peipei Li, Shuhan Xia, Shengyang Liu, Zekun Li, Ran He
arXiv AI
4d ago

Can Multimodal Large Language Models Generate and Detect Multimodal Social Media Fake News?

The paper investigates whether multimodal large language models (MLLMs) can generate and detect realistic multimodal fake news on social media. Using a multi‑agent framework—comprising a story agent, an image agent, and a critic agent—the authors produced over 9,000 paired multimodal news posts across science, health, and entertainment domains. They benchmarked 16 open‑ and closed‑source MLLMs for automated detection and found that most models fall far short of human accuracy, especially in identifying image authenticity, highlighting the need for stronger defenses against social media fake news.

By Jiyao Yang, Yang Liu, Zhenyue Qin, Qingyu Chen, Xiuzhen Zhang
arXiv Computer Vision
Sep 22

Dissecting Agentic Forensics: The Role of Triage, Prompting, and Evidence Arbitration in Open-World Fake Image Detection

The paper investigates an agentic framework for open‑world fake image detection that combines specialist detectors with per‑detector triage, prompting, and conflict‑aware evidence arbitration. Experiments across six configurations and three multimodal large language model backbones reveal that naive detector fusion yields high false‑positive rates, while triage and prompting consistently filter unreliable evidence. The most significant improvement comes from the reasoning component: a stronger judge markedly outperforms a weaker one, especially under distribution shift, and overall manipulation recall is nearly saturated, highlighting that the key challenge lies in calibrating trust and arbitrating conflicting forensic evidence rather than detecting manipulations themselves.

By Xianlong Li (IMT School for Advanced Studies Lucca, Italy), Pietro Bongini (University of Siena, Italy), Niccol\'o Pancino (University of Siena, Italy), Marco Blanchini (IMT School for Advanced Studies Lucca, Italy), Benedetta Tondi (University of Siena, Italy), Mauro Barni (University of Siena, Italy)
arXiv AI
6d ago

What Improves Multimodal Misinformation Detection? Answers from a Large-Scale Empirical Study

The paper investigates how different multimodal design choices affect the performance of misinformation detection systems. Using over 3,375 experiments across three benchmark datasets and various pre‑trained vision and language models, the authors systematically compare design options and conduct robustness analyses. The study offers practical guidance on which choices improve detection, when they may fail silently, and which pipeline components most influence model behavior, addressing four key research questions.

By Akshit Sharma, Prashant W. Patil
arXiv AI
Aug 24

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

Vis-Poison is a novel attack that poisons multimodal retrieval-augmented generation systems by inserting attacker-controlled images as visual evidence, without altering any textual metadata. The attack uses an automated multi-agent approach to create visually plausible poisoned images and has been tested on two multimodal RAG pipelines, four embedding models, and six generation models. In black-box settings, Vis-Poison achieves an end-to-end success rate between 40.16% and 65.40% against 30,000-entry knowledge bases, and remains effective against various multimodal large language models with an average success rate above 60%.

By Rujin Liang, Zhongpu Chen, Yuhao Lei, Xin Miao