arXiv Computer Vision

A Study of the Limits of Collaborative DCT-Based Image Denoising via Interpretable Neural Networks

The paper introduces DeepBM3D, a fully differentiable neural network that emulates the collaborative filtering strategy of BM3D for image denoising. It integrates non‑local patch grouping, DCT‑domain filtering with learned Wiener weights, and multi‑stage refinement, guided by lightweight convolutional feature extractors. Experiments demonstrate that DeepBM3D outperforms classical and hybrid baselines, competes with FFDNet at low to moderate noise levels, and excels on images with repetitive textures.

arXiv AI
Jul 31

PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging

arXiv:2602. 21987v3 Announce Type: replace-cross Abstract: Low-dose CT images are essential for reducing radiation exposure in cancer screening, pediatric imaging, and longitudinal monitoring protocols, but their quality is often degraded by noise from low-dose acquisition, patient motion, or scanner limitations, affecting both clinical interpretation and downstream analysis.

By Jitindra Fartiyal, Pedro Freire, Sergei K. Turitsyn, Sergei G. Solovski
arXiv Machine Learning
Sep 16

A deep dictionary network-based foundation model for ultra-low-dose CT denoising

The paper introduces a deep dictionary network (DDN) foundation model designed for ultra‑low‑dose CT (ULDCT) denoising across multiple organs. By cascading convolutional sparse coding layers with iterative soft‑thresholding, the architecture offers inherent interpretability, while dynamic dictionary and threshold modules enhance representation. The model is pre‑trained on over one million normal‑dose CT images and fine‑tuned on multi‑organ ULDCT datasets, achieving state‑of‑the‑art performance that consistently outperforms existing methods.

By Baoshun Shi, Shuangyi Yang, Ke Jiang, Bin Zhu, Zhanli Hu, Huazhu Fu
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv Computer Vision
Aug 25

Pixel-Space Diffusion via Observation Operators

Pixel‑Space Diffusion via Observation Operators introduces a new framework for pixel‑space diffusion models that addresses a scale‑time mismatch in existing methods. By replacing fixed full‑image supervision with a time‑indexed observation trajectory that progresses from coarse structures to the full image, the model aligns supervision with the natural recovery order of image details. The approach employs Gaussian‑Lanczos operators and a GL‑CoDA decoder to refine features progressively, resulting in faster convergence and higher generation quality, achieving an FID of 1.52 on ImageNet‑256.

By Shaojie Guo, Lichen Ma, Haoyang Tong, Yu He, Zipeng Guo, Xiaoan Liu, Feng Yan, Yu Guo, Fei Wang, Junshi Huang, Yan Wang