arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
By Dipankar Sarkar
arXiv:2606. 27396v1 Announce Type: cross Abstract: Test-input generation for tensor kernels is folkloric.
By Dipankar Sarkar
The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv:2608. 12700v1 Announce Type: new Abstract: Systems that generate GPU kernels with language models report high correctness rates.
By Rishi Shah, Rishav Shrestha
The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.
By Teng-Ruei Chen
Apple Silicon machines with large unified memory allow large tensors on a desktop GPU, but PyTorch’s Metal Performance Shaders (MPS) backend silently returns incorrect results for batched matrix multiplication when the output exceeds $2^{32}$ elements. The authors swept over dtypes, memory layouts, shapes, and batch sizes around $2^{31}$ and $2^{32}$ elements, finding that errors arise when operands are transposed views or when contiguous inputs exceed $2^{32}$ elements, with the backward pass also affected. They reproduced the issue on multiple machines and macOS versions, confirmed correct behavior on an NVIDIA A100, and demonstrated real‑world impact in a sentiment classifier, releasing a harness and guard to prevent such silent failures.
By Junichiro Niimi
arXiv:2609.22991v1 Announce Type: cross
Abstract: Apple Silicon machines with 192GB or more of unified memory make it routine to place tensors with more than $2^{32}$ elements on a desktop GPU. We sh...
By Junichiro Niimi
AutoTuneBench introduces a trustworthy measurement protocol for evaluating how large language model agents auto‑tune GPU kernels and serving engines. The benchmark addresses four failure modes—strawman baselines, machine‑dependent timing, saturated tasks, and infrastructure defects—by enforcing code‑frozen protocols, database validation, anti‑cheat checks, pre‑registered comparisons, and external result anchoring. Using this protocol, the authors demonstrate that previously reported speedups are inflated, revealing more modest improvements across different engines and machines.
By Li Chen
arXiv:2606. 10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data.
By Aarav Bedi
arXiv:2609. 19611v1 Announce Type: cross Abstract: Tensor programs, as used in deep learning models, are a prime target for optimization, as small performance improvements can have a large impact across training or inference workloads.
By Paul Biberstein, Joseph Devietti, Mayur Naik
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
By Yunxiang Zhang (Xiangjun), Ping Yu (Xiangjun), Jianyu Wang (Xiangjun), Max (Xiangjun), Fan, Julian Reed, Azalia Mirhoseini, Will Su
arXiv:2609.21058v1 Announce Type: cross
Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
By Gaurav Agarwal, Ashish Garg, Isha Singhal