arXiv Machine Learning By Dipankar Sarkar

Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels

Read the original on arXiv Machine Learning →

arXiv:2607. 16228v1 Announce Type: new Abstract: Most tensor-kernel correctness tests go through a fixed-shape all close-style check with hand-picked absolute and relative tolerances.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.

By Mingzhe Du, Anh Tuan Luu, Dong Huang, See-Kiong Ng
arXiv Machine Learning
Sep 2

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.

By Teng-Ruei Chen
arXiv Machine Learning
Sep 24

Silent Failures Beyond the 32-Bit Index Range: A Differential Characterization of Large-Tensor Matrix Multiplication in PyTorch's MPS Backend

Apple Silicon machines with large unified memory allow large tensors on a desktop GPU, but PyTorch’s Metal Performance Shaders (MPS) backend silently returns incorrect results for batched matrix multiplication when the output exceeds $2^{32}$ elements. The authors swept over dtypes, memory layouts, shapes, and batch sizes around $2^{31}$ and $2^{32}$ elements, finding that errors arise when operands are transposed views or when contiguous inputs exceed $2^{32}$ elements, with the backward pass also affected. They reproduced the issue on multiple machines and macOS versions, confirmed correct behavior on an NVIDIA A100, and demonstrated real‑world impact in a sentiment classifier, releasing a harness and guard to prevent such silent failures.

By Junichiro Niimi