arXiv Machine Learning

Spatially Adaptive Noise Injection

Spatially Adaptive Noise Injection (SANI) is a new diffusion sampling framework that adjusts the amount of noise added at each pixel during reverse diffusion. Unlike traditional samplers that apply a uniform noise variance across the image, SANI uses a probabilistic gating mechanism to inject noise only where the denoiser is uncertain, such as edges and textures, while preserving smooth regions. Experiments show that SANI consistently improves Fréchet Inception Distance over vanilla DDPM and DDIM samplers across various timesteps, and remains competitive with variance‑learning baselines.

arXiv Machine Learning
Sep 17

Beyond Random Couplings: Contrastive Noise Alignment in Generative Flows

The paper introduces Contrastive Noise Alignment (CNA), a training-time method for generative flow models that dynamically aligns Gaussian noise with data samples using a cross-modal InfoNCE objective. By modeling noise as an interacting particle system and regularizing with angular entropy and radial norm penalties, CNA reduces arbitrary data-noise couplings and flow curvature. Empirical results show that CNA improves generation quality, lowering FID by over 50% for few-step pixel-space generation compared to standard rectified flow and outperforming optimal transport baselines by at least 24%.

By Lennart Wittke, Vinicius Azevedo
arXiv Computer Vision
Sep 3

Local Epistemic Uncertainty Guided Active Sampling for Plug-and-play Diffusive Image Restoration

The paper introduces LEADer, a framework that uses local epistemic uncertainty to guide active sampling in diffusion-based image restoration. By adjusting prior strength per pixel and pruning sampling trajectories based on uncertainty traces, LEADer balances detail preservation with artifact suppression and accelerates convergence. The method is plug‑and‑play, theoretically guarantees data consistency and stable convergence, and improves performance across multiple state‑of‑the‑art diffusion models with minimal memory overhead.

By Jiaqi Zhang, Zheng Pang, Rongrong Gao, Qiyuan Zhang, Yang Yang
arXiv Computer Vision
Aug 25

Pixel-Space Diffusion via Observation Operators

Pixel‑Space Diffusion via Observation Operators introduces a new framework for pixel‑space diffusion models that addresses a scale‑time mismatch in existing methods. By replacing fixed full‑image supervision with a time‑indexed observation trajectory that progresses from coarse structures to the full image, the model aligns supervision with the natural recovery order of image details. The approach employs Gaussian‑Lanczos operators and a GL‑CoDA decoder to refine features progressively, resulting in faster convergence and higher generation quality, achieving an FID of 1.52 on ImageNet‑256.

By Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang
Hugging Face Trending Papers
Aug 13

PixSDS: Why Latent SDS Makes Noisy Pixels

Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high-frequency texture noise. We identify a failure mode of latent SDS caused by VAE-induced pixel drift: the optimized image can move along pixel-space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts.

arXiv AI
4d ago

NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training

arXiv:2605.08144v2 Announce Type: replace-cross Abstract: Training a diffusion model involves two sources of randomness for each data sample: the timestep and the Gaussian noise realization. The time...

By Haokai Zhao, Da Xing, Hanqun Cao, Tinson Xu, Xinyu Xiang, Yanchao Li, Xiangru Tang, Hongbin Lin, Zehong Wang, Kuan Pang, Peng Xia, Molei Tao, Li Erran Li, Aditya Joshi, Jure Leskovec, Fang Wu