arXiv Machine Learning

Accelerating Diffusion Transformers with Gaussian Process Rectified Feature Cache

The paper introduces GP-Refiner, a plug‑and‑play framework that uses Gaussian Process Regression to correct feature predictions in Diffusion Transformers. By observing that residuals between cached and full‑compute features follow a local zero‑mean Gaussian distribution, the method treats cached features as noisy observations of the true trajectory, enabling online correction without needing explicit labels. Experiments show that integrating GP‑Refiner with existing acceleration techniques, such as TaylorSeer, cuts computational load by 19.3% while improving image quality metrics (PSNR +0.9 dB, LPIPS 0.46→0.29).

arXiv AI
Aug 3

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

arXiv:2605. 04215v3 Announce Type: replace-cross Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over the traditional autoregressive paradigm.

By Michael Rottoli, Subhankar Roy, Stefano Paraboschi
arXiv Machine Learning
Aug 7

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.

By Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie