The paper introduces Spectral Correction Guidance, a training‑free method that uses spectral alignment to detect and correct deviations in guided diffusion trajectories. By comparing intermediate states to an analytic reference spectrum, the approach improves consistency with the forward process and enhances image generation quality. Experiments show consistent gains over baseline guidance in text‑to‑image tasks and on ImageNet, with benefits across guidance scales and fewer denoising steps.
By Gihoon Kim, Taesup Kim
Steering Fields introduce adaptive vector fields that re-estimate steering directions at each step of a flow-based text-to-image generation process, replacing the fixed global steering vectors traditionally used. By operating on noisy states, they provide a continuous trade-off between steering strength and content preservation, and allow simultaneous induction and inhibition of concepts without explicit spatial masks or object priors. The method achieves state-of-the-art safety steering benchmarks and can also function as a structure-preserving image-editing technique, delivering high semantic fidelity while remaining model-agnostic and inversion-free.
By Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer
arXiv:2510.03075v4 Announce Type: replace-cross
Abstract: Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models....
By Karim Farid, Rajat Sahay, Yumna Ali Alnaggar, Simon Schrodi, Volker Fischer, Cordelia Schmid, Thomas Brox
The paper introduces Global Transport (GT), a class‑agnostic optimal‑transport coupling that can be computed without class labels. GT associates different conditions with distinct regions of the source noise, which degrades performance when used without guidance but consistently improves generation when combined with classifier‑free guidance across various domains, model scales, and sampling budgets. The authors argue that coupling design should be evaluated under guided flow conditions rather than unguided generation, and demonstrate GT’s benefits on both discrete class‑conditioned and continuous text‑conditioned image generation.
By Katarina Petrovi\'c, Zander W. Blasingame, Danyal Rehman, \.Ismail \.Ilkan Ceylan, Michael Bronstein, Stephen Y. Zhang, Lazar Atanackovic, Alexander Tong
arXiv:2605. 31162v1 Announce Type: cross Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored.
By Shreyansh Modi, Akshat Tomar, Aarush Aggarwal
arXiv:2602. 20360v2 Announce Type: replace Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used in their vanilla conditional form: in image generation, samples without guidance often appear diffuse and lack fine-grained detail.
By Runlong Liao, Jian Yu, Baiyu Su, Chi Zhang, Lizhang Chen, Qiang Liu
arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.
By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv:2608.29233v1 Announce Type: new
Abstract: We present the winning solution to the ACM Multimedia 2026 Grand Challenge on Single-Image Guided Multi-Angle Image Synthesis. It ranks first among 293...
By Jie Li, Xingchen Zou, Yuxuan Liang
RefDiT is a new framework for reference-guided image generation that addresses the shortcomings of previous methods when handling complex scenes with multiple objects. It introduces local region guidance by decomposing a single identifier token into attribute-level signals, allowing the model to learn correspondences between tokens and specific regions of a reference image. The approach incorporates a low-rank adapter (LoRA) within a diffusion transformer (DiT) to adjust the inference prompt based on user-provided guidance context, thereby enabling more precise local attribute control.
By Rameshwar Mishra, Srikrishna Karanam, A V Subramanyam
arXiv:2606. 28696v1 Announce Type: new Abstract: Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation.
By Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan
Image outpainting extends an image beyond its original borders, requiring seamless style integration and globally coherent scene completion. Building on the success of diffusion models, recent methods have achieved substantial improvements in visual quality.
arXiv:2609.37492v1 Announce Type: new
Abstract: Instruction-guided image editing should change what the instruction names and leave the rest of the image untouched. In dual classifier-free guidance (...
By Zeyan Li, Wei Zhou, Hadi Amirpour, Minghao Zou, Panqi Yang, Jianfeng Xu