arXiv AI By Eugene Hauptmann, Nataliya Kosmyna

RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Machine Learning
Aug 4

Nova: An End-to-End MLIR Compiler for Deep Learning

arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.

By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv Machine Learning
Sep 21

Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

The paper reports on programming AMD XDNA NPUs for the FlashAttention workload using open‑source IRON and MLIR‑AIR compiler tools. It compares four reference designs on XDNA 1 and XDNA 2, showing that a fused kernel that keeps QKᵀ scores in local memory achieves 3.62 TFLOP/s on XDNA 2, doubling throughput and greatly improving energy efficiency over the IRON design and the integrated GPU. Roofline analysis guides when to fuse or stream operators based on each device’s ridge points, and the authors release the reference designs as open source.

By Erwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber, Samuel Bayliss