Apple Silicon machines with large unified memory allow large tensors on a desktop GPU, but PyTorch’s Metal Performance Shaders (MPS) backend silently returns incorrect results for batched matrix multiplication when the output exceeds $2^{32}$ elements. The authors swept over dtypes, memory layouts, shapes, and batch sizes around $2^{31}$ and $2^{32}$ elements, finding that errors arise when operands are transposed views or when contiguous inputs exceed $2^{32}$ elements, with the backward pass also affected. They reproduced the issue on multiple machines and macOS versions, confirmed correct behavior on an NVIDIA A100, and demonstrated real‑world impact in a sentiment classifier, releasing a harness and guard to prevent such silent failures.
By Junichiro Niimi
arXiv:2607. 14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.
By A. C. Opus, J. Q. Lu
arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.
By Dipankar Sarkar
The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim
The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.
By Teng-Ruei Chen
arXiv:2608. 10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.
By Matt J. Borowski, Blazej Osinski
The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv:2609.21058v1 Announce Type: cross
Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
By Gaurav Agarwal, Ashish Garg, Isha Singhal
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
By Om Mohite
arXiv:2608. 13057v1 Announce Type: cross Abstract: In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU.
By Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song
arXiv:2607. 19456v1 Announce Type: cross Abstract: We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly from the forward-pass Denotational Normal Form (DNF) of with the query-row index fixed to the current decode step.
By Lenore Mulin, Gaetan Hains
arXiv:2608.24738v1 Announce Type: new
Abstract: Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e....
By Kai Zhao