π€ Kernels: Major Updates
Related stories
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels
Easily Build and Share ROCm Kernels with Hugging Face
KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
arXiv:2607. 27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task.
kAgent: An execution-guided crash resolution agent for the Linux kernel
arXiv:2504. 20412v3 Announce Type: replace-cross Abstract: Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive.
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization
arXiv:2608. 09160v1 Announce Type: new Abstract: Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs).
Function calling and other API updates
Weβre announcing updates including more steerable API models, function calling capabilities, longer context, and lower prices.
Creating custom kernels for the AMD MI300
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Kernel Faults
arXiv:2505. 19489v2 Announce Type: replace Abstract: The Linux kernel is a critical system, serving as the foundation for numerous systems.
Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
arXiv:2607. 14541v1 Announce Type: new Abstract: Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads.
Tile-Level Activation Overlap for Efficient LLM Inference
arXiv:2607. 02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.