Diffusion and generative media

Image, video and audio generation — diffusion models, flow matching and the systems built on top of them.

5,061 stories · RSS feed

arXiv Machine Learning
1d ago

RW-Flow: One-Step Generation on Compact Manifolds via Wasserstein Gradient Flows

RW-Flow presents a new one‑step generative framework for data on compact manifolds, leveraging Wasserstein gradient flows. The authors derive a necessary and sufficient identifiability condition for velocity fields on compact, connected Riemannian manifolds, showing that a symmetric, Lipschitz‑continuous cost function yields identifiability iff its Gibbs kernel is nondegenerate. Experiments on geospatial events, protein and RNA torsion angles, and discretized manifolds demonstrate that RW‑Flow surpasses existing one‑step methods across most benchmark settings.

By Ualibyek Nurgulan, Seungwoo Yoo, Prin Phunyaphibarn, Minhyuk Sung
arXiv Machine Learning
1d ago

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.

By Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan
arXiv Computer Vision
1d ago

LensBridge: Frequency-Guided Compound Degradation Adaptation for Lens Aberration Correction and Veiling Glare Removal

LensBridge is a two‑stage framework that extends reusable aberration‑correction models to handle both lens aberrations and veiling glare (VG). Stage I builds a PSF‑aware diffusion foundation using discrete degradation priors from a large Lens Library, enabling aberration correction without explicit PSF input. Stage II adapts this foundation to compound degradation via frequency‑guided techniques—Frequency‑guided Degradation Completion (FDC) synthesizes training pairs and Frequency‑guided Pseudo Decomposition (FPD) conditions separate adaptation branches—allowing joint aberration correction and VG removal with only a few unpaired target observations.

By Xiaolong Qian, Zhonghua Yi, Qi Jiang, Kailun Yang, Shuhang Xie, Shaohua Gao, Kaiwei Wang