arXiv:2608. 12876v1 Announce Type: cross Abstract: Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit three failure modes from their training data: real and fake images collected from different sources invite provenance shortcuts, supervised explanation corpora teach templated rationales, and a static forgery corpus leaves the decision boundary standing still while generators keep moving.
By Yicheng Bao, Xiahui Guo, Xuhong Wang, Xin Tan
The paper investigates how visual understanding and generation objectives interact within unified multimodal models (UMMs). At the representation level, each objective enriches the other, but forcing them through the same computation path can cause one to dominate; a task‑decoupled architecture mitigates this. At the task and system levels, the authors demonstrate bidirectional transfer between shared knowledge and superior performance of an end‑to‑end UMM over a planner–executor pipeline on complex tasks.
By Penghao Wu, Haiwen Diao, Weichen Fan, Lewei Lu, Dahua Lin, Ziwei Liu
UniEvo‑VL is a self‑evolving framework that lets multimodal models improve themselves by using their own critiques as privileged information. The method trains a single model to act as both teacher and student, minimizing divergence between their diffusion distributions over sampling trajectories. Experiments on Qwen‑image‑2512 show significant gains in image generation metrics, and stronger external critics further raise the improvement ceiling, though results vary across tasks.
By Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi
SAVOR is a training framework for multimodal large language models that adds token and answer confidence to the output schema, optimises a Group Relative Policy Optimisation objective to penalise calibration error and poor abstention, and uses the learned confidence at inference to revisit visual evidence only when uncertain. Experiments on POPE, HallusionBench, AMBER, and MMHal-Bench with InternVL3-8B and Qwen3-VL-8B backbones show that SAVOR reduces hallucination while maintaining general capability on MME and MMBench, achieving lower Expected Calibration Error than DPO and decoding baselines.
By Zixiu Ding, Zilin Zhao, Yingjie He, Xinlang Kang, Guansu Wang, Wei Zhang
arXiv:2608.24372v1 Announce Type: new
Abstract: AI-generated image quality assessment (AIGIQA) requires jointly reasoning about perceptual fidelity and prompt alignment, two quality dimensions that a...
By Baoliang Chen, Qing Lin, Sijie Mai
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that...