RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.
arXiv:2608. 01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries.
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
The paper reports on programming AMD XDNA NPUs for the FlashAttention workload using open‑source IRON and MLIR‑AIR compiler tools. It compares four reference designs on XDNA 1 and XDNA 2, showing that a fused kernel that keeps QKᵀ scores in local memory achieves 3.62 TFLOP/s on XDNA 2, doubling throughput and greatly improving energy efficiency over the IRON design and the integrated GPU. Roofline analysis guides when to fuse or stream operators based on each device’s ridge points, and the authors release the reference designs as open source.
arXiv:2609.13612v1 Announce Type: new Abstract: Modern AI systems are built on the Transformer architecture, whose core operation, attention, accounts for the majority of computation and memory cost....