The paper introduces MATE, a reinforcement‑learning‑based post‑training framework for unified multimodal models that lets the generation and understanding branches challenge each other instead of cooperating. In MATE, each branch proposes candidate outputs that the other must reproduce, and the solver is trained on the worst‑handled candidate, creating an evolving adversarial loop without a separate adversary. Experiments on Janus‑Pro‑1B show that MATE improves generation and understanding metrics, including GenEval (+2.4), DPG‑Bench (+1.7), and an average of nine understanding benchmarks (+0.7), while enhancing consistency across image‑text cycles.
By Wentao Zhou, Weijie Gan, Jiayun Wang
arXiv:2410. 01574v4 Announce Type: replace-cross Abstract: The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse.
By Sina Mavali, Jonas Ricker, David Pape, Asja Fischer, Lea Sch\"onherr
arXiv:2609.14316v1 Announce Type: new
Abstract: Advances in image generation have made synthetic images increasingly difficult to distinguish from real photographs, raising concerns about the trustwo...
By Manni Cui, Ruiqi Liu, Zijian Yu, Hao Tan, Zibo Wei, Zian Wang, Ziheng Qin, Huijia Zhu, Weiqiang Wang, Jun Lan, Shu Wu
arXiv:2503. 01734v3 Announce Type: replace-cross Abstract: Attacks on machine learning models have been extensively studied through stateless optimization.
By Kyle Domico, Jean-Charles Noirot Ferrand, Ryan Sheatsley, Eric Pauley, Josiah Hanna, Patrick McDaniel
The paper investigates an agentic framework for open‑world fake image detection that combines specialist detectors with per‑detector triage, prompting, and conflict‑aware evidence arbitration. Experiments across six configurations and three multimodal large language model backbones reveal that naive detector fusion yields high false‑positive rates, while triage and prompting consistently filter unreliable evidence. The most significant improvement comes from the reasoning component: a stronger judge markedly outperforms a weaker one, especially under distribution shift, and overall manipulation recall is nearly saturated, highlighting that the key challenge lies in calibrating trust and arbitrating conflicting forensic evidence rather than detecting manipulations themselves.
By Xianlong Li (IMT School for Advanced Studies Lucca, Italy), Pietro Bongini (University of Siena, Italy), Niccol\'o Pancino (University of Siena, Italy), Marco Blanchini (IMT School for Advanced Studies Lucca, Italy), Benedetta Tondi (University of Siena, Italy), Mauro Barni (University of Siena, Italy)
arXiv:2511. 04949v2 Announce Type: replace-cross Abstract: Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trust.
By Tharindu Fernando, Clinton Fookes, Sridha Sridharan