Hugging Face Trending Papers

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inversion--transforming images back to latent noise for faithful reconstruction and editing--remains a challenging bottleneck due to the discretization errors of linear solvers.

arXiv Computer Vision
Sep 4

FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow

FlashRender is a few-step generative rendering framework that quickly retakes a source video along a target camera trajectory. It addresses discretization error by introducing Representation Transformation and Alignment (RETA) to align source-video representations with target-video features, reducing denoising trajectory curvature. The model is further refined with a MeanFlow objective and on-policy flow map distillation, achieving video quality and geometric consistency comparable to multi-step baselines at a 25× lower sampling cost while improving camera controllability.

By Byeongjun Park, Byung-Hoon Kim, Hyungjin Chung
arXiv Computer Vision
Sep 16

Bi-FlowGS: Bridging Generative View Completion and Gaussian Geometry through Bidirectional Flow Co-Refinement

Bi-FlowGS introduces a bidirectional co-refinement framework that links generative view completion with 3D Gaussian Splatting geometry. It employs Video-to-Geometry Flow Distillation (V2G) to transfer temporal correspondence from restored videos into Gaussian geometry, mitigating the Geometry Cheating problem. Simultaneously, Geometry-to-Video Flow-Guided Restoration (G2V) uses the current 3DGS geometry to guide temporally consistent video restoration, creating a loop where restored videos and optimized geometry iteratively improve each other, leading to better rendering quality and geometric consistency on wide-baseline and 360° benchmarks.

By Yuetong Wang, Jinsheng Quan, Yi Yang, Yawei Luo
arXiv Machine Learning
Sep 10

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

arXiv:2609.08084v1 Announce Type: cross Abstract: Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computati...

By Igor Pavlovic, Thiemo Wandel, Anton Obukhov, Luca Bartolomei, Andrey Davydov, Fabio Tosi, Matteo Poggi, Sabine S\"usstrunk, Dengxin Dai
arXiv Computer Vision
Sep 21

Geometry-Aware Diffusion Guidance via Curvature-Adaptive Tubular Correction

The paper introduces Curvature-Adaptive Tubular Correction (CAT), a training‑free plugin that refines diffusion guidance by decomposing the guidance gradient into normal and tangent components and regulating them within a noise‑dependent geometric budget. CAT charges normal displacement at first order and tangent displacement according to directional curvature, solving a one‑dimensional dual equation for optimal magnitudes and using Armijo backtracking to calibrate the step size. Experiments on seven inverse problems with FFHQ and ImageNet demonstrate that CAT consistently improves pixel‑ and latent‑space samplers, enhances perceptual metrics, and achieves the lowest FID across classifier‑free guidance scales while maintaining stable saturation and contrast.

By Enze Jiang, Jinwei He, Zheng Ma
arXiv AI
Sep 1

EquiReg: Equivariance Regularized Diffusion for Inverse Problems

EquiReg introduces an equivariance‑regularized diffusion framework that penalises sampling trajectories deviating from the data manifold, thereby improving posterior sampling for inverse problems. By formalising manifold‑preferential equivariant functions—naturally arising from data augmentation or inherent symmetries—EquiReg guides diffusion steps toward symmetry‑preserving regions of the solution space. The method shows consistent gains in both linear and nonlinear image restoration tasks and partial differential equation solving, especially under reduced sampling and measurement consistency steps, and is available as open‑source code.

By Bahareh Tolooshams, Aditi Chandrashekar, Rayhan Zirvi, Abbas Mammadov, Jiachen Yao, Chuwei Wang, Anima Anandkumar