arXiv:2605. 18714v2 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture.
By Songsong Yu, Yuxin Chen, Ying Shan, Yanwei Li
arXiv:2607. 07117v1 Announce Type: cross Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image.
By Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
By Zhihong Chen, Xuehai Bai, Yang Shi, Chaoyou Fu, Huanyu Zhang, Haotian Wang, Xiaoyan Sun, Zhang Zhang, Liang Wang, Yuanxing Zhang, Pengfei Wan, Yi-Fan Zhang
The paper introduces a capability‑centric data infrastructure for generalist image generation, integrating task‑specific supervision with a curriculum that aligns with the dependencies among generative capabilities. It employs three interoperable data engines—text‑image grounding, inter‑image transformation, and image‑knowledge association—alongside caption experts to harmonize text‑to‑image and editing supervision. The system curates massive corpora (440M T2I images, 120M editing pairs, 27M image‑entity pairs) and trains multimodal diffusion models (3B and 6B parameters) from scratch, achieving broad visual coverage and versatile rendering as shown by CPI‑Bench and qualitative tests.
By Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.
By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick
WeAgent-MMGenEdit is a comprehensive framework for multimodal agentic image generation and editing that addresses the unreliability of current models when prompts require external world knowledge. It introduces a multimodal harness with persistent evidence management, a scalable data construction pipeline producing 23K supervised trajectories and 14.7K RL tasks, and a bilingual benchmark (WeBench-MMGenEdit) for knowledge-intensive generation and multi-image editing. Post‑training methods based on SFT and RL further refine the agent policy and image backend, enabling a 30B‑parameter policy to outperform similarly sized models and approach the performance of a 1T‑parameter agent.
By Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng