Should We Skip Diffusion?
arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
arXiv:2607. 05319v1 Announce Type: cross Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures.
arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides fea...
arXiv:2411.05005v2 Announce Type: replace-cross Abstract: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, m...
arXiv:2512. 20963v3 Announce Type: replace Abstract: Diffusion models excel at generating high-quality, diverse samples, yet they risk memorizing training data when overfit to the training objective.
arXiv:2609.24919v1 Announce Type: new Abstract: Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag...
Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions.
The paper introduces D3S2, a diffusion‑guided dataset distillation framework tailored for semantic segmentation. It tackles long‑tailed class imbalance, pixel‑wise alignment, and high computational cost by first selecting class‑balanced masks and then synthesizing images with a pretrained diffusion model conditioned on those masks. Guided diffusion sampling further refines the data with segmentation‑consistency and class‑wise feature matching losses, achieving superior performance at a 1% compression rate on ADE20K and COCO‑Stuff.
arXiv:2606. 31699v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that isolated features can serve as controllable intervention points.
arXiv:2510. 17136v2 Announce Type: replace Abstract: The generation of high-quality, diverse, and prompt-aligned images is a central goal in image-generating diffusion models.
The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.
arXiv:2607. 15693v1 Announce Type: cross Abstract: We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be readily understood from its parameters.
arXiv:2608. 01298v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.
arXiv:2607. 06856v1 Announce Type: cross Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics.