arXiv AI

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

arXiv:2607. 06114v1 Announce Type: cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs).

arXiv AI
Sep 25

Accelerating Video Diffusion via Training-Free Trajectory Routing

The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.

By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh
arXiv Computation and Language
Sep 23

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

The paper introduces PACE-dLLM, an acceleration method for diffusion language models (dLLMs) that uses the model’s own per‑step confidence to estimate a ‘confidence cliff’ and determine the optimal look‑ahead horizon for block decoding. By fitting this cliff in closed form at each step, PACE-dLLM sets the horizon to its saturation point and applies an independent confidence threshold for token commitment, thereby avoiding the trade‑offs inherent in fixed‑size block decoding. Experiments on reasoning and code benchmarks show that PACE-dLLM achieves the best average accuracy on open‑source dLLM backbones while delivering significant wall‑clock speedups—up to 5.23× on LLaDA and 3.06× on Dream—improving the quality‑throughput Pareto frontier.

By Xiaocheng Lu, Shuhan Guo, Ziyue Ma, Jie Zhang, Jian Liu, Jingcai Guo, Haoxuan Che, Song Guo
arXiv Computer Vision
Sep 3

SelfLift: Accelerating Few-Step Diffusion via Self-Recovering Resolution Transition

SelfLift is a progressive‑resolution framework that accelerates few‑step diffusion models by enabling late, self‑recovering transitions between low‑ and high‑resolution latents. It introduces a training‑free Artifact‑Aware Consistency Lift that uses disagreement between direct latent lifting and pixel‑VAE re‑encoding to detect and correct artifacts, and a self‑recovery policy that transfers high‑resolution guidance from an internal teacher. Experiments on FLUX.2‑Klein and Z‑Image‑Turbo show latency reductions of 41.5% and 44.1%, and overall speedups of 29.61× and 19.21× over 50‑step baselines while maintaining competitive generation quality.

By Tingyan Wen, Chenqian Yan, Xurui Peng, Xiazhang Fang, Shuai Wang, Xueqian Wang, Songwei Liu
Hugging Face Trending Papers
Sep 24

Accelerating Video Diffusion via Training-Free Trajectory Routing

Accelerating Video Diffusion via Training-Free Trajectory Routing (TRACK) introduces a heterogeneous denoising strategy that switches between large and small diffusion models at selected steps, determined by a calibration process that measures disagreement between model predictions. By routing quality-sensitive steps to the large model and low-disagreement steps to the small model, TRACK achieves significant speedups—up to 2.73×—across several video diffusion benchmarks while maintaining comparable quality and diversity. The method requires no retraining, architectural changes, or online dual-model evaluation, making it a practical acceleration paradigm for video diffusion.

arXiv Machine Learning
Jun 30

Momentum Guidance: Plug-and-Play Guidance for Flow Models

arXiv:2602. 20360v2 Announce Type: replace Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used in their vanilla conditional form: in image generation, samples without guidance often appear diffuse and lack fine-grained detail.

By Runlong Liao, Jian Yu, Baiyu Su, Chi Zhang, Lizhang Chen, Qiang Liu
arXiv AI
Sep 3

GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories

GeoSPRINT is a training‑free framework that constructs non‑uniform sampling schedules for diffusion model inference by detecting geometrically redundant steps in denoising trajectories. It uses a hyperplanarity test in latent space, implemented via QR factorization, to allocate more steps to high‑curvature regions, and introduces the trajectory projection score α_traj as a model‑free diagnostic for flow quality. Across CIFAR‑10, LSUN Church, and Stable Diffusion v1.5, GeoSPRINT consistently outperforms uniform DDIM schedules at matched NFE budgets, improving FID scores by up to 1.93 points.

By Arpita Joshi
Hugging Face Trending Papers
Aug 6

LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction

Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollouts for policy exploration and optimization. Existing GRPO methods for flow models therefore replace the inference-time ODE with a stochastic differential equation (SDE) during training.