arXiv Computer Vision By Zeyang Liu, Le Wang, Sanping Zhou, Yuxuan Wu, Xiaolong Sun, Gang Hua, Haoxiang Li

UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation

Read the original on arXiv Computer Vision →

UniLayDiff is a Unified Diffusion Transformer that tackles content‑aware layout generation across a wide range of tasks—unconditional, element‑type, size, and relationship‑conditioned generation—using a single end‑to‑end trainable model. It treats layout constraints as a distinct modality within a Multi‑Modal Diffusion Transformer framework, capturing interactions among background images, layout elements, and constraints. The model is further refined with LoRA fine‑tuning to incorporate relation constraints, achieving state‑of‑the‑art performance and, according to the authors, the first unified solution for all content‑aware layout generation tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 11

Learning Interaction between Image and Layout Priors for Joint Image-Layout Generation in Design Templates

The paper introduces InterIL, a joint generative model that simultaneously produces a background image and a layout of foreground elements for graphic design templates, addressing the limitations of sequential generation approaches. InterIL connects pretrained image and layout diffusion backbones via a learnable communication module, freezing the backbones to preserve prior knowledge while training only the interaction module. The model also offers a test‑time guidance strategy, enabling users to impose preferences without retraining, and demonstrates superior image, layout, and harmonization quality compared to previous methods.

By Shirong Yang, Bo Yang, Ying Cao
arXiv AI
Jul 21

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.

By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo