Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv Machine Learning
Sep 30

Mixture-of-Kittens: MoE Megakernel for NVL72s

arXiv:2609.36070v1 Announce Type: cross Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...

By Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher R\'e
arXiv AI
Sep 30

Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

The paper introduces Prediction‑Aligned Context Compaction (PACC), a method that learns a compact memory representation for long‑video generation by distilling a frozen video generator. PACC trains a compressor to aggregate past frames into memory tokens, using the generator as both teacher and student during on‑policy distillation. Experiments on MBench and VBench‑Long show that PACC improves memory‑event coverage and consistency, achieving better scores than strong baselines and producing competitive minute‑long videos.

By Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng, Lydia E. Kavraki, Zhiwei Steven Wu
arXiv Machine Learning
Sep 30

Improved Distributional Diffusion Models

arXiv:2609.37147v1 Announce Type: cross Abstract: Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a \emph{distributional} denoiser trained via a scoring rule...

By Tommaso Martorella, Alexandre Galashov, Felix Krause, Stefan Andreas Baumann, Valentin De Bortoli, Arthur Gretton, Bj\"orn Ommer
arXiv Machine Learning
Sep 30

On State Reduction in Linear Attention

arXiv:2602.04852v3 Announce Type: replace Abstract: Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate tha...

By Philipp Nazari, T. Konstantin Rusch