arXiv Machine Learning

Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

The paper reports on programming AMD XDNA NPUs for the FlashAttention workload using open‑source IRON and MLIR‑AIR compiler tools. It compares four reference designs on XDNA 1 and XDNA 2, showing that a fused kernel that keeps QKᵀ scores in local memory achieves 3.62 TFLOP/s on XDNA 2, doubling throughput and greatly improving energy efficiency over the IRON design and the integrated GPU. Roofline analysis guides when to fuse or stream operators based on each device’s ridge points, and the authors release the reference designs as open source.

arXiv AI
Jul 13

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

arXiv:2607. 09385v1 Announce Type: cross Abstract: The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs).

By Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini
arXiv AI
Jun 3

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

arXiv:2605. 25645v2 Announce Type: replace-cross Abstract: We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation.

By Jatin Kishnani, Mayank Goel, Amit Singh, Pulkit Agrawal, Sairanjan Mishra