arXiv AI

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation

arXiv:2607. 09133v1 Announce Type: cross Abstract: While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency.

arXiv Machine Learning
Jun 30

Momentum Guidance: Plug-and-Play Guidance for Flow Models

arXiv:2602. 20360v2 Announce Type: replace Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used in their vanilla conditional form: in image generation, samples without guidance often appear diffuse and lack fine-grained detail.

By Runlong Liao, Jian Yu, Baiyu Su, Chi Zhang, Lizhang Chen, Qiang Liu
arXiv AI
2d ago

Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond

Steering Fields introduce adaptive vector fields that re-estimate steering directions at each step of a flow-based text-to-image generation process, replacing the fixed global steering vectors traditionally used. By operating on noisy states, they provide a continuous trade-off between steering strength and content preservation, and allow simultaneous induction and inhibition of concepts without explicit spatial masks or object priors. The method achieves state-of-the-art safety steering benchmarks and can also function as a structure-preserving image-editing technique, delivering high semantic fidelity while remaining model-agnostic and inversion-free.

By Simone Facchiano, Jan Eric Lenssen, Bernt Schiele, Wolfgang Stammer, Fabio Galasso, Jonas Fischer
arXiv Computer Vision
Sep 22

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

arXiv:2609.23153v1 Announce Type: new Abstract: Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}:...

By Yuxi Liu, Haoyu Li, Zekun Zhang, Tengxu Sun, Yixiang Cai, Jiayong Li, Yifei Xia, Tianle Liu, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kai Zhang, Kun Yuan, Bin Cui
Hugging Face Trending Papers
Aug 6

Energy-Guided Flow Matching

Pixel-space generative models bypass lossy latent compression, yet necessitate joint learning of global structure and fine-grained details in a high-dimensional space. Standard flow matching interpolates noise toward a fixed clean-image endpoint, leaving the spectral evolution to be learned implicitly.