The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim
arXiv:2609.21058v1 Announce Type: cross
Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
By Gaurav Agarwal, Ashish Garg, Isha Singhal
arXiv:2604. 23466v2 Announce Type: replace Abstract: NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (TMA) efficiency on modern GPUs.
By Divakar Kumar Yadav, Tian Zhao, Deepak Kumar
arXiv:2606. 06510v1 Announce Type: cross Abstract: Conventional HPC dogma holds that native hardware FP64 silicon is the irreducible foundation of scientific computing -- the "holy grail" of double-precision simulation.
By Satoshi Matsuoka
arXiv:2609.14507v1 Announce Type: cross
Abstract: Single-GPU long-context inference with Mixture-of-Experts (MoE) models requires spilling the key-value cache (KVCache) to CPU memory. The spilled KV...
By Enda Yu, Dezun Dong, Xiangke Liao
arXiv:2606. 08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run.
By Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen, Thanh Q. Duong, Linh D. Le, Duy M. H. Nguyen, Vien A. Ngo, An T. Le
The article argues that on AI‑optimised NVIDIA B300 GPUs and newer, the FP8 tensor‑core matrix operation—implemented via the CRT‑based Ozaki Scheme II—can become the primary substrate for matrix‑heavy FP64 kernels while maintaining FP64‑grade accuracy. It introduces the Tensor‑Memory Equilibrium (TME) model, a Roofline extension with four parameters, to show that FP8 can match native FP64 performance under certain intensity thresholds and tile‑fusion conditions. The study identifies two notable exceptions—large dense‑square DGEMM and 3‑D FFT—where additional hardware or software adjustments are required to reach the memory roof.
whyItMatters":"The paper demonstrates that FP8, with appropriate reconstruction and deconstruction strategies, can replace native FP64 for high‑performance computing workloads on modern GPUs, potentially reducing hardware complexity and energy consumption while preserving accuracy."
By Satoshi Matsuoka
arXiv:2608. 10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.
By Matt J. Borowski, Blazej Osinski
arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.
By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
The paper introduces a multi‑shell decoder for Leech‑lattice vector quantization, achieving the best reported 2‑bit quality under its evaluation protocol. It presents a GPU‑friendly layout that fuses dequantization with matrix‑vector multiplication, demonstrating significant speed and memory advantages over traditional one‑hot masks and other 4‑bit methods. Experiments show the new kernel outperforms baseline approaches across multiple model sizes, with measurable gains in throughput and reduced byte traffic.
By Pier-Jean Malandrino (Scub)
arXiv:2608. 15383v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models reduce arithmetic by activating only a small subset of experts per token, yet deployment still requires storing and moving the full expert bank.
By Amjad Saab
The paper introduces Right In-Place (RiP) convolution, a memory‑efficient strategy that corrects and generalizes previous in‑place convolution formulations to arbitrary stride, dilation, padding, and rectangular kernels. RiP aligns each layer’s input and output within a shared workspace, enabling safe, row‑major access with minimal memory overhead. Experiments on 10,000 random layers and 84 layers from 25 architectures show no corruption, matching or improving on existing herringbone workspaces while reducing memory usage by up to 24.8% and lowering peak activation memory on Raspberry Pi Pico MCUs by 12.5–33.3% without affecting cycle counts.
By Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye Eigbe