arXiv AI

Anchor-Conditioned Compositional Control for Landscape Image Generation

arXiv:2606. 07638v1 Announce Type: cross Abstract: Image generative models, though widely used as creative tools, offer limited support for the kind of compositional control that photographers and visual artists routinely exercise.

arXiv AI
5d ago

Correcting Guided Diffusion Trajectories with Spectral Alignment

The paper introduces Spectral Correction Guidance, a training‑free method that uses spectral alignment to detect and correct deviations in guided diffusion trajectories. By comparing intermediate states to an analytic reference spectrum, the approach improves consistency with the forward process and enhances image generation quality. Experiments show consistent gains over baseline guidance in text‑to‑image tasks and on ImageNet, with benefits across guidance scales and fewer denoising steps.

By Gihoon Kim, Taesup Kim
arXiv AI
Oct 1

Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond

Steering Fields introduce adaptive vector fields that re-estimate steering directions at each step of a flow-based text-to-image generation process, replacing the fixed global steering vectors traditionally used. By operating on noisy states, they provide a continuous trade-off between steering strength and content preservation, and allow simultaneous induction and inhibition of concepts without explicit spatial masks or object priors. The method achieves state-of-the-art safety steering benchmarks and can also function as a structure-preserving image-editing technique, delivering high semantic fidelity while remaining model-agnostic and inversion-free.

By Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer
arXiv Machine Learning
3d ago

Global Transport Couplings for Classifier-Free Guided Flows

The paper introduces Global Transport (GT), a class‑agnostic optimal‑transport coupling that can be computed without class labels. GT associates different conditions with distinct regions of the source noise, which degrades performance when used without guidance but consistently improves generation when combined with classifier‑free guidance across various domains, model scales, and sampling budgets. The authors argue that coupling design should be evaluated under guided flow conditions rather than unguided generation, and demonstrate GT’s benefits on both discrete class‑conditioned and continuous text‑conditioned image generation.

By Katarina Petrovi\'c, Zander W. Blasingame, Danyal Rehman, \.Ismail \.Ilkan Ceylan, Michael Bronstein, Stephen Y. Zhang, Lazar Atanackovic, Alexander Tong
arXiv Machine Learning
Jun 30

Momentum Guidance: Plug-and-Play Guidance for Flow Models

arXiv:2602. 20360v2 Announce Type: replace Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used in their vanilla conditional form: in image generation, samples without guidance often appear diffuse and lack fine-grained detail.

By Runlong Liao, Jian Yu, Baiyu Su, Chi Zhang, Lizhang Chen, Qiang Liu
arXiv AI
Jul 21

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.

By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv Computer Vision
Sep 7

RefDiT: Local Attribute Guidance in Reference-Based Image Generation

RefDiT is a new framework for reference-guided image generation that addresses the shortcomings of previous methods when handling complex scenes with multiple objects. It introduces local region guidance by decomposing a single identifier token into attribute-level signals, allowing the model to learn correspondences between tokens and specific regions of a reference image. The approach incorporates a low-rank adapter (LoRA) within a diffusion transformer (DiT) to adjust the inference prompt based on user-provided guidance context, thereby enabling more precise local attribute control.

By Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam
arXiv AI
Jun 30

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

arXiv:2606. 28696v1 Announce Type: new Abstract: Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation.

By Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan