arXiv Machine Learning

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

Hugging Face Trending Papers
Jun 29

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions.

arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv Computer Vision
Sep 24

ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming

ZoomDiff is a high‑fidelity diffusion model designed to improve dual‑camera smooth zooming by producing photo‑realistic transitions. It strengthens dual‑image conditional guidance during denoising, injects flow‑aligned multi‑scale features from the VAE encoder into the decoder to recover high‑frequency details, and uses flow‑guided temporal consistency supervision to ensure smoother transitions. Experiments on synthetic and real‑world datasets show that ZoomDiff outperforms state‑of‑the‑art methods both quantitatively and qualitatively.

By Jiayi Zhang, Renlong Wu, Yukang Ding, Sibin Deng, Wangmeng Zuo