arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
By Dipankar Sarkar
arXiv:2606. 27396v1 Announce Type: cross Abstract: Test-input generation for tensor kernels is folkloric.
By Dipankar Sarkar
The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv:2608. 12700v1 Announce Type: new Abstract: Systems that generate GPU kernels with language models report high correctness rates.
By Rishi Shah, Rishav Shrestha
The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.
By Teng-Ruei Chen
Apple Silicon machines with large unified memory allow large tensors on a desktop GPU, but PyTorch’s Metal Performance Shaders (MPS) backend silently returns incorrect results for batched matrix multiplication when the output exceeds $2^{32}$ elements. The authors swept over dtypes, memory layouts, shapes, and batch sizes around $2^{31}$ and $2^{32}$ elements, finding that errors arise when operands are transposed views or when contiguous inputs exceed $2^{32}$ elements, with the backward pass also affected. They reproduced the issue on multiple machines and macOS versions, confirmed correct behavior on an NVIDIA A100, and demonstrated real‑world impact in a sentiment classifier, releasing a harness and guard to prevent such silent failures.
By Junichiro Niimi