Creating custom kernels for the AMD MI300
Related stories
🤗 Kernels: Major Updates
Introducing the AMD 5th Gen EPYC™ CPU
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels
Hugging Face on AMD Instinct MI300 GPU
AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification
AsmEvo is an agentic assembly-level optimizer that targets compiled AMDGPU code objects, reconstructing a reassemblable representation and applying low-level edits guided by a long-horizon agent. It rebuilds ABI-preserving optimized objects and verifies functional equivalence through differential testing against the original binary. Experiments show significant speedups—up to 1.35× geometric mean on MI308X and 1.18× on MI300X—while maintaining correctness.
Easily Build and Share ROCm Kernels with Hugging Face
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
Fine-Tune MMS Adapter Models for low-resource ASR
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
arXiv:2607. 20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads.
Towards Automated Kernel Generation in the Era of LLMs
arXiv:2601. 15727v3 Announce Type: replace Abstract: The performance of modern AI systems is fundamentally constrained by the quality of their underlying GPU kernels, which translate high-level algorithmic semantics into low-level hardware operations.
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.