arXiv Machine Learning

Patch-PODiff-ViT: Structured Latent Diffusion with Patchwise POD for Super-Resolution and Uncertainty Quantification

arXiv:2606. 31290v1 Announce Type: new Abstract: Diffusion models enable probabilistic super-resolution and conditional generation, but pixel-space methods are computationally expensive and learned latent spaces often lack interpretable uncertainty quantification.

arXiv Computer Vision
Aug 27

Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

UGDiff introduces an uncertainty-guided diffusion paradigm for single-image super-resolution, aiming to improve the perception‑distortion trade‑off. The method estimates reconstruction uncertainty of latent features from a high‑fidelity image and uses this uncertainty, along with diffusion sampler posterior variance, to selectively restore high‑frequency details in uncertain regions while preserving fidelity elsewhere. Experiments show that UGDiff outperforms state‑of‑the‑art diffusion‑based SR methods.

By Ren Wang, Yung-Yu Chuang
arXiv Machine Learning
Sep 3

Perceptually Regularized Diffusion Model for Image Super-Resolution

The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.

By Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang, Min Wang, Jing Qin, Yifei Lou, Weihong Guo
arXiv AI
Sep 10

Diffusion Model in Latent Space for Medical Image Segmentation Task

The paper introduces MedSegLatDiff, a diffusion-based framework that combines a variational autoencoder (VAE) with a latent diffusion model for medical image segmentation. By compressing images into a low-dimensional latent space, the method reduces noise and speeds up training, while a weighted cross‑entropy loss preserves tiny structures such as small nodules. Evaluated on ISIC‑2018, CVC‑Clinic, and LIDC‑IDRI datasets, MedSegLatDiff achieves state‑of‑the‑art Dice and IoU scores, generates diverse segmentation hypotheses, and produces confidence maps that enhance interpretability and reliability for clinical deployment.

By Ngoc Huynh Trinh, Hai Toan Nguyen, Son Ba Luong, Quoc Long Tran
Hugging Face Trending Papers
Aug 5

Intrinsic-Hybrid Latent Diffusion Models for Generative Modeling on Unknown Manifolds

We introduce the Intrinsic Hybrid Latent Diffusion Model (ILDM), a generative framework that integrates probabilistic dimensionality reduction with geometry-aware diffusion on unknown manifolds. While diffusion models (DMs) have achieved state-of-the-art results in high-dimensional data synthesis, they rely on large training datasets and ignore intrinsic geometric structure.

arXiv AI
Aug 3

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

arXiv:2607. 29337v1 Announce Type: cross Abstract: Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of annotated data in medical imaging, particularly in optical coherence tomography (OCT) of mouse eyes, where manual retinal layer delineation is labour-intensive due to tiny structures and required expertise, resulting in scarce datasets.

By Fernando Garc\'ia-Torres, Roc\'io del Amor, Sandra Morales, \'Alvaro Barroso, Peter Heiduschka, Bj\"orn Kemper, Valery Naranjo
arXiv Computer Vision
Sep 15

MedDiME: Efficient Latent Diffusion with Adaptive Masking for Medical Counterfactual Generation

MedDiME is a latent-space, classifier‑guided diffusion framework designed for medical counterfactual image generation. It introduces a gradient‑driven adaptive masking mechanism that works directly in latent space, enabling spatially precise edits while avoiding the high computational and memory costs of pixel‑space methods. Experiments show MedDiME can produce high‑quality counterfactuals up to 40× faster and using 13× less GPU memory than previous diffusion baselines.

By Yan Zeng, Changlu Guo, Anders Nymark Christensen, Morten Rieger Hannemose, Anders Bjorholm Dahl