Should We Skip Diffusion?
arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
The paper introduces a decoder that generates images whose features closely match a user-specified feature, enabling detailed analysis of a vision-related deep neural network’s feature space. Implemented as a guided diffusion model, it steers a pre-trained diffusion model to minimize the Euclidean distance between the feature of a clean image and the target feature at each generation step. The method is training‑free, works on a single COTS GPU, and has been validated on CLIP’s image encoder and ResNet‑50, showing high feature‑matching accuracy and practical feasibility.
arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model. Most recent methods adapt large pretrained text-to-image (T2I) latent diffusion models for their strong capacity and generative priors.
arXiv:2608. 01298v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.
arXiv:2607. 17411v1 Announce Type: cross Abstract: Inverse rendering is traditionally solved via differentiable renderers and gradient descent, which requires substantial problem-specific engineering and is prone to getting stuck in local minima due to ambiguities.
arXiv:2607. 06982v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks.
arXiv:2403. 07711v5 Announce Type: replace-cross Abstract: Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extending these models to video generation.
arXiv:2606. 00094v1 Announce Type: cross Abstract: Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, low-dimensional, and compact parameterization space.
arXiv:2609.24919v1 Announce Type: new Abstract: Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag...
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
arXiv:2607. 06856v1 Announce Type: cross Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics.
Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resource-constrained embedded devices.
V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.