arXiv AI

GuidedBridge: Training-freely Improving Bridge Models with Prior Guidance

arXiv:2606. 03119v1 Announce Type: cross Abstract: Guidance methods, such as classifier-free guidance (CFG) and auto-guidance (AG), have advanced noise-to-data generation in diffusion models.

Hugging Face Trending Papers
Jun 29

UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

Recent advances in diffusion models have shown impressive performance in controllable image generation and dense prediction tasks. However, existing approaches typically treat diffusion-based controllable generation and dense prediction as separate tasks, overlooking the potential benefits of jointly modeling the heterogeneous distributions.

arXiv Machine Learning
1d ago

Learned End-to-End Guidance Schedules for Diffusion Models

The paper introduces Learned End-to-End Guidance Schedules (LEEGS) for diffusion models, which train a time‑dependent guidance schedule to balance data quality and requirement satisfaction while reducing sampling steps. LEEGS minimizes the guidance function over a small set of examples using stochastic gradient descent and employs a gradient approximation to cut training time by a factor of four. Experiments on tasks such as image inpainting, noisy image inverse problems, face‑ID‑guided generation, and PDE problems show that LEEGS outperforms baselines at the same computational budget or matches constant guidance with only 10% of the steps.

By Aneesh Barthakur, Mathias Niepert, Luiz F. O. Chamon
arXiv AI
Sep 2

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

V-Co investigates visual co-denoising for pixel-space diffusion models, using a unified JiT-based framework to isolate key design choices. The study identifies two essential components: a dual-stream architecture with flexible cross-stream interaction and a perceptual-drifting hybrid loss combined with RMS-based feature rescaling for stronger semantic supervision. Experiments on ImageNet-256 demonstrate that V-Co surpasses baseline pixel-space diffusion and strong prior pixel-diffusion methods at comparable model sizes while requiring fewer training epochs.

By Han Lin, Xichen Pan, Zun Wang, Yue Zhang, Chu Wang, Jaemin Cho, Mohit Bansal
arXiv Machine Learning
Aug 27

WAVE: Reversing the Guidance Hierarchy for Coarse-to-Fine Guided Depth Super-Resolution

WAVE introduces a multi-level discrete wavelet transform (ML‑DWT) to reverse the typical fine‑to‑coarse bias in guided depth super‑resolution. By consuming wavelet sub‑bands and semantic tokens in reverse order, it separates structure and detail reconstruction, applies semantic gating to high‑frequency bands, and fuses modalities via an invertible coupling mechanism. Experiments on multiple benchmarks show that WAVE matches or outperforms existing methods, especially at high upsampling factors where low‑resolution depth has minimal structure.

By Tayyab Nasir, Daochang Liu, Ajmal Mian
arXiv AI
6d ago

Neural Bridge Processes

Neural Bridge Processes (NBPs) replace the input‑independent forward kernel of Neural Diffusion Processes with an input‑anchored bridge trajectory, allowing conditioning inputs to influence the noisy training states. When input and output dimensions differ, NBPs learn an output‑space anchor that guides the generative path without altering the denoising backbone. Theoretical analysis shows that this anchoring yields pathwise input distinguishability, injects input information into noisy states, and provides a direct gradient pathway, leading to consistent performance gains across synthetic regression, EEG, CylinderFlow, and image regression tasks.

By Jian Xu, Yican Liu, Delu Zeng, John Paisley, Qibin Zhao