arXiv Computer Vision By Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng, Zhe Lin, Ming-Hsuan Yang, Truong Nguyen

Adversarial Training for Pixel Diffusion

Read the original on arXiv Computer Vision →

Pixel diffusion models generate RGB images directly but tend to miss fine‑scale natural‑image statistics. The authors introduce an adversarial post‑training step that adds an adversarial loss to the model’s output at non‑high‑noise timesteps, without changing the architecture or sampling procedure. This approach improves distribution fidelity, coverage, prompt alignment, and perceptual quality across two pixel backbones, and restores missing high‑frequency spectral power while avoiding memorization or mode dropping.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Aug 27

Continuous Adversarial Flow Models

The paper introduces continuous adversarial flow models, a continuous-time flow framework trained with an adversarial objective that replaces the fixed mean-squared-error criterion of flow matching. By incorporating a learned discriminator, the method guides training toward a different generalized distribution, yielding samples more closely aligned with the target data distribution. Applied as a post‑training step, it markedly improves ImageNet 256px generation metrics—reducing the guidance‑free FID of latent‑space SiT from 8.26 to 3.63 and of pixel‑space JiT from 7.17 to 3.57—and also enhances guided generation and text‑to‑image benchmarks.

By Shanchuan Lin, Ceyuan Yang, Zhijie Lin, Hao Chen, Haoqi Fan