arXiv AI By Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

Read the original on arXiv AI →

TOLA is a diffusion‑based text image super‑resolution method that eliminates iterative image‑text diffusion by using a one‑step latent adaptation framework. It employs a confidence‑weighted text conditioning module to build a reliable semantic condition and a lightweight latent residual correction module to fix structured residual errors, thereby preserving text fidelity. Experiments show TOLA outperforms existing diffusion‑based TSR methods, achieving at least 2.72 dB higher PSNR on the CTR‑TSR‑Test benchmark.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

The paper introduces a 3D-CLIP encoder trained with structured hard negatives to improve vision‑language alignment for text‑to‑CT generation. This encoder drives a latent diffusion model that operates directly in 3D latent space, eliminating spatial artifacts from super‑resolution pipelines. Experiments on the CT‑RATE dataset show state‑of‑the‑art image fidelity and factual correctness across 18 pathological conditions, with lower inference time and GPU memory usage than competing methods.

By Daniele Molino, Camillo Maria Caruso, Filippo Ruffini, Paolo Soda, Valerio Guarrasi
arXiv Computer Vision
Aug 28

Zero-Shot Video Restoration and Enhancement with Text-to-Image Latent Diffusion Models and Multi-Modal References

The paper introduces a zero‑shot video restoration and enhancement framework that leverages a text‑to‑image latent diffusion model along with multi‑modal references. It employs dual prompt tuning inversion and sampling to cut inference time to about one‑third of the original, while also strengthening performance and temporal consistency. Additional techniques such as texture‑aware video token merging, referenced self‑attention, and referenced token merging further improve temporal coherence across frames.

By Cong Cao, Huanjing Yue, Xin Liu, Jingyu Yang