ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming
Read the original on arXiv Computer Vision →ZoomDiff is a high‑fidelity diffusion model designed to improve dual‑camera smooth zooming by producing photo‑realistic transitions. It strengthens dual‑image conditional guidance during denoising, injects flow‑aligned multi‑scale features from the VAE encoder into the decoder to recover high‑frequency details, and uses flow‑guided temporal consistency supervision to ensure smoother transitions. Experiments on synthetic and real‑world datasets show that ZoomDiff outperforms state‑of‑the‑art methods both quantitatively and qualitatively.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.