arXiv AI

Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

arXiv:2509. 12046v2 Announce Type: replace-cross Abstract: Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned generation remains challenging due to the sparse nature of layout conditions and the risk of feature entanglement.

arXiv Computer Vision
Sep 11

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

The paper introduces InterIL, a joint generative model that simultaneously produces a background image and a layout of foreground elements for graphic design templates, addressing the limitations of sequential generation approaches. InterIL connects pretrained image and layout diffusion backbones via a learnable communication module, freezing the backbones to preserve prior knowledge while training only the interaction module. The model also offers a test‑time guidance strategy, enabling users to impose preferences without retraining, and demonstrates superior image, layout, and harmonization quality compared to previous methods.

By Shirong Yang, Bo Yang, Ying Cao
Hugging Face Trending Papers
Jun 11

InterleaveThinker: Reinforcing Agentic Interleaved Generation

Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation.

arXiv AI
Jul 21

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.

By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv AI
3d ago

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

The paper presents a post‑training approach for text‑to‑image models that combines a preference reward, trained on large human preference data, with rubric‑based rewards that assess prompt faithfulness and other desirable traits. The authors show that a simple reward composition strategy outperforms a naive weighted average, leading to significant Elo gains on the Arena leaderboard for models like Flux2dev and Ideogram‑4. They also release Arena‑T2I‑Training, a 1K subset of data to aid reproducible research in post‑training.

By Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh
arXiv Computer Vision
Sep 22

Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation

arXiv:2609.22916v1 Announce Type: new Abstract: Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit la...

By Guanqiao Chen, Jingru Tan, Dongxing Mao, Catherine Chen, Zijian Du, Libo Qin, Hu Jian Guo, Alex Jinpeng Wang
arXiv AI
Aug 7

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

arXiv:2603. 08652v2 Announce Type: replace Abstract: Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning.

By Haodong Li, Chunmei Qing, Huanyu Zhang, Dongzhi Jiang, Yihang Zou, Hongbo Peng, Dingming Li, Yuhong Dai, ZePeng Lin, Juanxi Tian, Yi Zhou, Siqi Dai, Jingwei Wu, Pheng-Ann Heng