Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts
arXiv:2608. 01740v1 Announce Type: new Abstract: Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps.
Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters.
arXiv:2608. 01740v1 Announce Type: new Abstract: Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps.
arXiv:2603.01623v2 Announce Type: replace Abstract: Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference spe...
arXiv:2607. 27842v1 Announce Type: cross Abstract: Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive.
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
arXiv:2607. 29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising.
Accelerating Video Diffusion via Training-Free Trajectory Routing (TRACK) introduces a heterogeneous denoising strategy that switches between large and small diffusion models at selected steps, determined by a calibration process that measures disagreement between model predictions. By routing quality-sensitive steps to the large model and low-disagreement steps to the small model, TRACK achieves significant speedups—up to 2.73×—across several video diffusion benchmarks while maintaining comparable quality and diversity. The method requires no retraining, architectural changes, or online dual-model evaluation, making it a practical acceleration paradigm for video diffusion.
The paper introduces Physics‑SIMS‑TS, a conditional diffusion model designed for long‑horizon oil and gas production forecasting. It enforces monotone decline through negative guidance, decline‑curve constraints, and isotonic projection during sampling, and incorporates spatial training augmentation and an ensembled stochastic sampler to produce calibrated predictive distributions. Evaluated on over 35,000 wells across three jurisdictions, Physics‑SIMS‑TS achieves the highest accuracy among diffusion forecasters and matches transformer ensembles, with only a 0.5% increase in mean squared error for monotonicity.
arXiv:2606. 27688v1 Announce Type: cross Abstract: In financial forecasting, predictive performance depends not only on which model is trained, but also on how the trained model is deployed.
arXiv:2606. 04342v1 Announce Type: cross Abstract: Multi-step time series forecasting (MSF) is commonly evaluated using point-wise error metrics such as mean squared error (MSE), implicitly treating the conditional mean as a sufficient target.
arXiv:2606. 26778v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs.
arXiv:2608. 11235v1 Announce Type: new Abstract: Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon.
The paper introduces Internal Dual-Wiener routing (Internal‑DW), a backward‑only method that weight‑balances internal gradient routes in autoregressive forecasting. By estimating bounded Wiener gains for identity and nonlinear paths, it suppresses unpredictable noise while preserving predictable learning signals, reducing forecast error by 5.2%–13.8% on four weak‑drive testbeds compared to full BPTT and outperforming gradient clipping, Jacobian regularization, and truncated BPTT in most cases. The approach shows that long‑horizon supervision can be effective without trusting every backward gradient equally.