arXiv Machine Learning

Should We Skip Diffusion?

arXiv Computer Vision
Oct 1

Looped Diffusion Transformer

arXiv:2609.40305v1 Announce Type: new Abstract: Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternat...

By Yong Xien Chng, Tianyi Chen, Wenwen Tong, Haiwen Diao, Zhongang Cai, Lei Yang, Ziwei Liu, Lewei Lu, Dahua Lin, Gao Huang
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv Computer Vision
1d ago

Feature Space Analysis by Guided Diffusion Model

The paper introduces a decoder that generates images whose features closely match a user-specified feature, enabling detailed analysis of a vision-related deep neural network’s feature space. Implemented as a guided diffusion model, it steers a pre-trained diffusion model to minimize the Euclidean distance between the feature of a clean image and the target feature at each generation step. The method is training‑free, works on a single COTS GPU, and has been validated on CLIP’s image encoder and ResNet‑50, showing high feature‑matching accuracy and practical feasibility.

By Kimiaki Shirahama, Kaduki Yamashita, Miki Yanobu, Miho Ohsaki