arXiv AI

M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification

arXiv Machine Learning
Jun 11

MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs

arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.

By Xinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji, Jinchen Jiang, Zepeng Zhao, Ziruo Xiao, Zihao Ye, Yingyi Huang, Ruihang Lai, Hongyi Jin, Bohan Hou, Mengdi Wu, Yixin Dong, Anthony Yip, Zihao Ye, Songting Wang, Wenqin Yang, Xupeng Miao, Tianqi Chen, Zhihao Jia
arXiv AI
Jul 29

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.

By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim
arXiv AI
Sep 2

CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language

CUDA‑Harness is a framework that enables the generation and optimization of CUDA kernels directly from natural language. It introduces Intermediate‑Structured Generation to bridge high‑level semantics with low‑level kernel code, uses Synthesis‑Based Verification to mitigate reward hacking by providing isolated test data, and employs Feedback‑Adaptive Evolution to prioritize correctness while improving performance. Experiments show the approach generalizes across different large language models, hardware platforms, and even supports C‑to‑CUDA transpilation.

By Qi Fan, An Zou, Yehan Ma
arXiv Machine Learning
Aug 4

TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference

arXiv:2608. 01975v1 Announce Type: cross Abstract: Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the inference engine, Python/C++ backend, host CUDA APIs, GPU kernels, and distributed communication.

By Ruilin Xu, Junyi Li, Pengfei Chen, Zongxuan Xie
arXiv AI
4d ago

AI as a Compiler: Compiling Triton kernels without the Triton compiler

The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.

By Fran\c{c}ois Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini