arXiv AI

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

TOLA is a diffusion‑based text image super‑resolution method that eliminates iterative image‑text diffusion by using a one‑step latent adaptation framework. It employs a confidence‑weighted text conditioning module to build a reliable semantic condition and a lightweight latent residual correction module to fix structured residual errors, thereby preserving text fidelity. Experiments show TOLA outperforms existing diffusion‑based TSR methods, achieving at least 2.72 dB higher PSNR on the CTR‑TSR‑Test benchmark.

arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv Computer Vision
Aug 28

Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

The paper introduces a zero‑shot video restoration and enhancement framework that leverages a text‑to‑image latent diffusion model along with multi‑modal references. It employs dual prompt tuning inversion and sampling to cut inference time to about one‑third of the original, while also strengthening performance and temporal consistency. Additional techniques such as texture‑aware video token merging, referenced self‑attention, and referenced token merging further improve temporal coherence across frames.

By Cong Cao, Huanjing Yue, Xin Liu, Jingyu Yang
arXiv AI
Sep 16

Efficient Text-to-Image Generation: An Adaptive Step Schedule Controller for Diffusion Models

The paper introduces an adaptive step schedule controller for text‑to‑image diffusion models, allowing the number of denoising steps to vary based on the complexity of the input prompt. By mixing step schedules of different sizes and monitoring error discrepancies at each timestep, the method switches schedules to maintain image quality while reducing inference time. Experiments on COCO and DiffusionDB demonstrate that this approach achieves faster generation without sacrificing visual fidelity.

By Kuluhan Binici, Cihan Acar, Shivam Aggarwal, Siying Liu, Tulika Mitra
arXiv Computer Vision
Sep 3

GlyphAnchor: Enhancing Visual Text Rendering via Position-Anchored Glyph Priors

GlyphAnchor is a new method that improves visual text rendering in image generation and editing models by adding lightweight glyph patch conditions anchored to the target image’s positional encoding. The approach is trained with staged supervised finetuning and text-aware post‑training, and it works with both text‑to‑image and image‑editing diffusion transformers. Experiments on various backbones and the newly introduced InfoTextBench benchmark show that GlyphAnchor consistently enhances text fidelity while maintaining overall image quality, especially for long, complex, or densely arranged text and rare characters.

By Qiang Xiang, Shuang Sun, Binglei Li, Yibo Chen, Xu Tang, Yao Hu, Junping Zhang
arXiv Computer Vision
Aug 27

Uncertainty-Guided Latent Diffusion Models for Faithful Super Resolution

UGDiff introduces an uncertainty-guided diffusion paradigm for single-image super-resolution, aiming to improve the perception‑distortion trade‑off. The method estimates reconstruction uncertainty of latent features from a high‑fidelity image and uses this uncertainty, along with diffusion sampler posterior variance, to selectively restore high‑frequency details in uncertain regions while preserving fidelity elsewhere. Experiments show that UGDiff outperforms state‑of‑the‑art diffusion‑based SR methods.

By Ren Wang, Yung-Yu Chuang