arXiv AI

How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?

arXiv Machine Learning
Sep 23

Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.

By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim
arXiv AI
Jul 29

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.

By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
arXiv AI
Jun 30

KernelSight-LM: A Kernel-Level LLM Inference Simulator

arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.

By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt
arXiv AI
Jul 10

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

arXiv:2605. 10886v3 Announce Type: replace-cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8.

By Liang Luo, Yinbin Ma, Quanyu Zhu, Vasiliy Kuznetsov, Yuxin Chen, Neng Shi, Jian Jiao, Jiecao Yu, Buyun Zhang, Tongyi Tang, Xiaohan Wei, Yanli Zhao, Zeliang Chen, Yuchen Hao, Venkatesh Ranganathan, Sandeep Parab, Yantao Yao, Maxim Naumov, Chunzhi Yang, Shen Li, Ellie Wen, Wenlin Chen, Santanu Kolay, Chunqiang Tang
Hugging Face Trending Papers
Sep 24

KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

KernelOPT is a multi-agent system that optimizes GPU kernels generated by compilers like PyTorch Inductor by treating compiled models as structured artifacts. It preserves vendor library calls and focuses on Triton sub-kernels, using five profiling-guided LLM agents and a four-gate verification cascade to filter and validate candidates. On 250 KernelBench problems, KernelOPT achieves geometric mean speedups of 1.40×, 1.15×, and 1.07× over torch.compile at different optimization levels.

arXiv AI
Sep 25

KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization

KernelOPT is a multi‑agent system that optimizes GPU kernels generated by compilers like PyTorch Inductor by treating compiled models as structured artifacts. It preserves vendor library calls and focuses on Triton sub‑kernels, using five profiling‑guided LLM agents and a four‑gate verification cascade to ensure correctness and performance before re‑stitching the model. On 250 KernelBench problems, KernelOPT achieves geometric mean speedups of 1.40×, 1.15×, and 1.07× over torch.compile at three optimization levels.

By Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur