arXiv Computer Vision

MIAR: Medical Image Super-Resolution With Autoregressive Modeling

MIAR introduces a multi‑scale autoregressive framework for medical image super‑resolution, treating the task as a conditional, progressive next‑scale prediction. It incorporates a Scale‑Adaptive Structural Decoder to preserve structural fidelity and uses a hierarchical beam search during inference to reduce recursive error accumulation. Experiments show MIAR outperforms existing methods, achieving a 7.86% MUSIQ improvement and a 2.02× speedup over diffusion‑based approaches.

arXiv Machine Learning
Sep 3

Perceptually Regularized Diffusion Model for Image Super-Resolution

The paper introduces a perceptually regularized diffusion framework for image super‑resolution, adding perceptual‑loss based regularization to the standard diffusion training objective. This approach incorporates prior knowledge to improve training convergence and encourages the recovery of meaningful image features. Experiments on benchmark datasets show enhanced perceptual quality while maintaining competitive distortion metrics.

By Chuxiangbo Wang, Pavithra Venkatachalapathy, Ying Liang, Min Wang, Jing Qin, Yifei Lou, Weihong Guo
arXiv Computer Vision
Sep 2

Prior-Guided Implicit Neural Representations for Single-Subject Diffusion MRI Super-Resolution

The paper introduces a transfer‑learning framework that pre‑trains an implicit neural representation (INR) on a high‑resolution diffusion MRI template and then adapts it to individual subjects through registration and fine‑tuning. This approach enables native single‑subject super‑resolution, achieving a 4× through‑plane up‑sampling from 5 mm to 1.25 mm on Human Connectome Project data. Compared to a recent baseline, the method reduces NRMSE by 36–49 % and increases FSIM by 24–43 %, while training 6× faster and outperforming other INR‑based techniques on both image quality and domain‑specific metrics.

By Abdulkader Ghandoura, Marsil Zakour, William Consagra, Yogesh Rathi
arXiv Computer Vision
Aug 27

Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

UGDiff introduces an uncertainty-guided diffusion paradigm for single-image super-resolution, aiming to improve the perception‑distortion trade‑off. The method estimates reconstruction uncertainty of latent features from a high‑fidelity image and uses this uncertainty, along with diffusion sampler posterior variance, to selectively restore high‑frequency details in uncertain regions while preserving fidelity elsewhere. Experiments show that UGDiff outperforms state‑of‑the‑art diffusion‑based SR methods.

By Ren Wang, Yung-Yu Chuang
arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah