Stable and Near-Reversible Diffusion ODE Solvers for Image Editing
arXiv:2605. 16399v2 Announce Type: replace-cross Abstract: The inversion of diffusion models plays a central role in image editing.
arXiv:2605. 16399v2 Announce Type: replace-cross Abstract: The inversion of diffusion models plays a central role in image editing.
Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inversion--transforming images back to latent noise for faithful reconstruction and editing--remains a challenging bottleneck due to the discretization errors of linear solvers.
arXiv:2601. 19180v2 Announce Type: replace-cross Abstract: Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines.
arXiv:2505. 06668v2 Announce Type: replace-cross Abstract: We present StableMotion, a novel framework that leverages geometric and content priors from pretrained large-scale image diffusion models for motion estimation in single-image rectification tasks such as Stitched Image Rectangling (SIR) and Rolling Shutter Correction (RSC).
arXiv:2605. 31162v1 Announce Type: cross Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored.
The paper introduces Curvature-Adaptive Tubular Correction (CAT), a training‑free plugin that refines diffusion guidance by decomposing the guidance gradient into normal and tangent components and regulating them within a noise‑dependent geometric budget. CAT charges normal displacement at first order and tangent displacement according to directional curvature, solving a one‑dimensional dual equation for optimal magnitudes and using Armijo backtracking to calibrate the step size. Experiments on seven inverse problems with FFHQ and ImageNet demonstrate that CAT consistently improves pixel‑ and latent‑space samplers, enhances perceptual metrics, and achieves the lowest FID across classifier‑free guidance scales while maintaining stable saturation and contrast.
arXiv:2606. 09901v1 Announce Type: cross Abstract: Diffusion-based generative models enable powerful image editing capabilities, but achieving precise control while maintaining fidelity and safety remains challenging.
arXiv:2607. 19895v1 Announce Type: cross Abstract: Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion.
Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent source-image conditioning can limit how fully an edit is executed and how natural the result appears, especially when the target scene diverges substantially from the input.
The paper introduces CAT-Flow, a pair of lightweight, training‑free algorithms—CAT‑OV and CAT‑OT—that adapt step‑sizes during Flow Matching inference by estimating curvature in time or state space. These methods avoid extra neural evaluations and achieve constant‑order truncation error bounds. Experiments show that CAT‑OV and CAT‑OT improve image quality metrics across four text‑to‑image Flow Matching models, cutting the required generation steps by up to 40%.
Spectral Feedback is a new algorithm for aligning discrete diffusion models at test time by iteratively revisiting and editing token positions rather than only steering the reverse process. It selects edit-sets—groups of token positions to re-mask and re-sample—using sparse Fourier representations of edit-set value functions, enabling efficient optimization of which tokens to revisit. The method is model-agnostic and improves alignment performance across pretrained, test‑time aligned, and fine‑tuned diffusion models, achieving significant gains in protein stability for inverse folding tasks.
arXiv:2604. 27147v3 Announce Type: replace-cross Abstract: In generative modeling, we often wish to produce samples that maximize a user-specified reward such as aesthetic quality or alignment with human preferences, a problem known as \textit{guidance}.