arXiv Machine Learning

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

arXiv:2607. 11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference.

arXiv Computation and Language
Sep 23

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Flash-dLLM is a training‑free inference acceleration framework that improves the speed and memory efficiency of Diffusion Large Language Models (dLLMs). It tackles GPU memory I/O bottlenecks by introducing an I/O‑aware fused KV‑cache kernel and then employs a draft‑and‑verify decoding strategy that uses the dLLM itself as both drafter and verifier. Experiments on mathematical reasoning and code‑generation tasks show Flash‑dLLM outperforms existing acceleration methods, achieving up to 11.0× speedups over the Elastic‑Cache baseline.

By Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen
arXiv AI
Jul 13

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

arXiv:2607. 09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs).

By Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini
arXiv Computation and Language
6d ago

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification

UniPrefill is a prefill‑acceleration framework that works with virtually any model architecture, directly speeding up token‑level computation. It is implemented as a continuous‑batching operator and extends vLLM’s scheduling to support prefill‑decode co‑processing and tensor parallelism. The method delivers up to a 2.1× speedup in Time‑to‑First‑Token, with gains growing as concurrent requests increase.

By Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He