UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.
By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick
arXiv:2607. 07117v1 Announce Type: cross Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image.
By Stepanida Alekseeva, Jenifer Kalafatovich, Seong-Whan Lee
While traditional graphics methods often synthesize 3D indoor scenes autoregressively or hierarchically, recent vision-language model (VLM)-based generators predominantly adopt a one-shot paradigm where the full layout is planned at once. This one-shot approach often requires global re-optimization or complete reconstruction during interactive editing (e.
arXiv:2608. 15006v1 Announce Type: cross Abstract: Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning.
By Penghao Yin, Haomin Wang, Qihong Tang, Xiaoye Qu, Hongjie Zhang, Xiao-Ping Zhang
SenseNova-U1.5 is an 8B‑MoT native unified multimodal model that can understand, reason about, and generate visual content without using an encoder or VAE. It improves visual fidelity and text rendering through spatially coherent patch reconstruction, large‑scale training on curated generation and editing data, and native resolutions up to 4K. Post‑training, specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing are optimized and distilled into a multi‑expert framework, yielding advances in image fidelity, complex composition, multi‑reference editing, and instruction following.
By Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng, Jiangnan Chen, Ruixi Zhang, Ruohui Wang, Wenwen Tong, Xiangyu Fan, Yubo Wang, Yue Zhu, Yuwei Niu, Zhengqi Bai, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Bo Yang, Chen Feng, Chengguang Lv, Guangjia Liu, Guanlin Wang, Hanyu Zhang, Haojia Yu, Hongcan Xiao, Hongli Wang, Huan Wu, Huaping Zhong, Jian Fang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jing Zuo, Jingcheng Ni, Junxiang Xu, Linjun Dai, Mutian Xu, Peishen Yan, Penghao Wu, Ruijie Mao, Ruisi Wang, Shihao Bai, Shuang Yang, Shuya Yang, Shuyan Zheng, Silei Wu, Siying Li, Tao Chu, Tianbo Zhong, Tongxi Zhou, Weichao Luo, Weichen Fan, Wenhao Jia, Wenjie Gao, Xiangli Kong, Yan Li, Yang Yong, Zimo Wen, Zixuan Qian, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation.
The paper introduces ReImaGin, a method that uses image generation models as a flexible visual reasoning tool for multimodal large language models. Unlike traditional fixed-function vision tools, ReImaGin accepts natural language commands and can perform open-ended visual operations such as removing occlusions or creating floorplans from multiple views. Experiments on six diverse visual reasoning tasks show that ReImaGin outperforms both text-only reasoning and specialist vision-tool baselines, achieving up to a 25% improvement.
By Nishad Singhi, Hector Garcia Rodriguez, Aditya Arora, Marcus Rohrbach, Anna Rohrbach
arXiv:2606. 04479v1 Announce Type: cross Abstract: Recent text-to-image (T2I) models can render highly legible and well-structured text within images, enabling applications including document generation and slide generation.
By Jiajun Hong, Jiawei Zhou
arXiv:2606. 24849v1 Announce Type: cross Abstract: Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved.
By Zixuan Li, Haokun Lin, Yicheng Xiao, Zhiwei Li, Xinyang Song, Zelong Zheng, Yong He, Heng Yao, Ke Ding, Chao Yu, Chuan Yuan, Qi Li, Zhenan Sun
CoVA‑SFT is a new large‑scale dataset comprising 51.9K samples and over 222K multimodal reasoning steps that teach models to interleave text and visual abstractions across five layout families and 17 complex tasks. It includes explicit rationale formulations, agentic renderings, and verification loops to help models build and maintain internal visual workspaces for purely textual reasoning problems. A companion benchmark, CoVA‑Bench, contains 1,700 held‑out test samples for reproducible evaluation, and models fine‑tuned on CoVA‑SFT outperform all interleaved CoT baselines by more than 2× on average, though they still lag behind strong text‑only CoT baselines.
By Tsung-Han Wu, Heekyung Lee, Anya Ji, Haoming Chen, Trevor Darrell, Joseph E. Gonzalez, David M. Chan
The paper introduces RIG-BENCH, a benchmark for evaluating reasoning-driven image generation (RIG) in four cognitively demanding domains—Concept-based, Transformation-based, Pattern & Structure, and Scenario-based—using 2000 curated samples. It highlights a reasoning-generation gap in current unified generative models (UGMs) and world simulators, noting that these models often produce locally plausible but globally illogical outputs. RIG-BENCH aims to serve as a rigorous stress test and diagnostic framework to guide the development of next-generation, logically grounded UGMs and simulators.
By Yutong Liu, Nan Huang, Xu Cao, James M. Rehg
The paper presents a systematic study of scientific image synthesis, comparing pixel‑based generation and programmatic approaches. It introduces ImgCoder, a logic‑driven framework that follows an "understand‑plan‑code" workflow to enhance structural precision, and SciGenBench, a benchmark that evaluates images for information utility and logical validity. The authors find that pixel‑based models exhibit systematic failure modes and that fine‑tuning large multimodal models on rigorously verified synthetic images consistently improves downstream reasoning performance.
By Honglin Lin, Zheng Liu, Chonghan Qin, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu