RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
arXiv:2608. 12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance.
The paper explores using large language models (LLMs) to replace traditional compiler backends, a process termed AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, achieving 0.83x–3.34x the performance of autotuned Triton on a variety of GPUs and ML kernels. The study also extends a PTX verifier to support modern GPU features, highlighting the potential for AI compilers to reduce engineering effort for new hardware.
arXiv:2608. 12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance.
arXiv:2606. 02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators.
arXiv:2604. 23466v2 Announce Type: replace Abstract: NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (TMA) efficiency on modern GPUs.
KernelOPT is a multi‑agent system that optimizes GPU kernels generated by compilers like PyTorch Inductor by treating compiled models as structured artifacts. It preserves vendor library calls and focuses on Triton sub‑kernels, using five profiling‑guided LLM agents and a four‑gate verification cascade to ensure correctness and performance before re‑stitching the model. On 250 KernelBench problems, KernelOPT achieves geometric mean speedups of 1.40×, 1.15×, and 1.07× over torch.compile at three optimization levels.
arXiv:2608. 12629v1 Announce Type: new Abstract: GPU kernel agents and GPU programming languages have advanced separately, leaving expert kernels difficult to reproduce.
AsmEvo is an agentic assembly-level optimizer that targets compiled AMDGPU code objects, reconstructing a reassemblable representation and applying low-level edits guided by a long-horizon agent. It rebuilds ABI-preserving optimized objects and verifies functional equivalence through differential testing against the original binary. Experiments show significant speedups—up to 1.35× geometric mean on MI308X and 1.18× on MI300X—while maintaining correctness.
The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
KernelOPT is a multi-agent system that optimizes GPU kernels generated by compilers like PyTorch Inductor by treating compiled models as structured artifacts. It preserves vendor library calls and focuses on Triton sub-kernels, using five profiling-guided LLM agents and a four-gate verification cascade to filter and validate candidates. On 250 KernelBench problems, KernelOPT achieves geometric mean speedups of 1.40×, 1.15×, and 1.07× over torch.compile at different optimization levels.
arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
arXiv:2604.18616v2 Announce Type: replace-cross Abstract: LLM coding agents can generate correct GPU kernels, but their performance still trails expert libraries. Reaching peak throughput requires co...
arXiv:2604. 01489v2 Announce Type: replace Abstract: High-performance GPU kernels are critical to modern machine learning systems, yet developing them remains a manual, expert-driven process.