ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
ZoomDiff is a high‑fidelity diffusion model designed to improve dual‑camera smooth zooming by producing photo‑realistic transitions. It strengthens dual‑image conditional guidance during denoising, injects flow‑aligned multi‑scale features from the VAE encoder into the decoder to recover high‑frequency details, and uses flow‑guided temporal consistency supervision to ensure smoother transitions. Experiments on synthetic and real‑world datasets show that ZoomDiff outperforms state‑of‑the‑art methods both quantitatively and qualitatively.
arXiv:2602.19202v3 Announce Type: replace Abstract: Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensit...
arXiv:2608. 05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame.
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
arXiv:2603. 17555v2 Announce Type: replace-cross Abstract: Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.
arXiv:2411.05005v2 Announce Type: replace-cross Abstract: Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, m...