Creating custom kernels for the AMD MI300
Related stories
๐ค Kernels: Major Updates
Introducing the AMD 5th Gen EPYCโข CPU
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels
Hugging Face on AMD Instinct MI300 GPU
Easily Build and Share ROCm Kernels with Hugging Face
Fine-Tune MMS Adapter Models for low-resource ASR
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
arXiv:2607. 20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads.
Towards Automated Kernel Generation in the Era of LLMs
arXiv:2601. 15727v3 Announce Type: replace Abstract: The performance of modern AI systems is fundamentally constrained by the quality of their underlying GPU kernels, which translate high-level algorithmic semantics into low-level hardware operations.
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
Kernel Foundry: A Diagnosis-driven Evolutionary Kernel Optimizer with Multi-Experts
arXiv:2605. 30359v2 Announce Type: replace-cross Abstract: Generating high-performance GPU kernels remains challenging due to the need for both correctness and hardware-aware optimization.
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
arXiv:2606. 11357v1 Announce Type: cross Abstract: With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets.