arXiv:2603.24595v2 Announce Type: replace-cross
Abstract: Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is...
By Mengting He, Shihao Xia, Haomin Jia, Wenfei Wu, Linhai Song
arXiv:2511. 10480v3 Announce Type: replace-cross Abstract: Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed workload execution.
By Changhai Man, Joongun Park, Hanjiang Wu, Huan Xu, Srinivas Sridharan, Tushar Krishna
arXiv:2606. 28565v1 Announce Type: cross Abstract: As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency targets.
By Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt
arXiv:2609.13585v1 Announce Type: cross
Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...
By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.
By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
arXiv:2607. 16241v1 Announce Type: cross Abstract: Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench.
By Yunxiang Zhang (Xiangjun), Ping Yu (Xiangjun), Jianyu Wang (Xiangjun), Max (Xiangjun), Fan, Julian Reed, Azalia Mirhoseini, Will Su
arXiv:2606. 02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators.
By Taras Sereda, Burak Bartan, Ankita Nayak, Tom St. John, Natalie Serrino, Zain Asgar
arXiv:2605.21603v1 Announce Type: cross
Abstract: Intra-device parallelism addresses resource under-utilization in ML inference and training by overlapping the execution of operators with different r...
By Yi Pan, Yile Gu, Jinbin Luo, Yibo Wu, Ziren Wang, Hongtao Zhang, Ziyi Xu, Shengkai Lin, Baris Kasikci, Stephanie Wang
arXiv:2604. 23466v2 Announce Type: replace Abstract: NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (TMA) efficiency on modern GPUs.
By Divakar Kumar Yadav, Tian Zhao, Deepak Kumar
arXiv:2607. 19539v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes.
By Minyu Cui, Anna Wingkvist, Morgan Ericsson
arXiv:2607. 10183v1 Announce Type: cross Abstract: Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory capacity, making offloading inference necessary to extend effective model capacity with CPU memory.
By Yangyijian Liu, Hongyi Ye, Mingyang Li, Wu-jun Li
The paper addresses the problem of non‑deterministic outputs from large language models (LLMs) when run on different GPU architectures, caused by floating‑point non‑associativity and hardware‑dependent kernel choices. It proposes a set of fixed‑configuration fused‑upcast GEMM kernels that load 16‑bit weights, upcast to FP32, and perform IEEE‑754 compliant reductions in a problem‑shape‑dependent order, ensuring identical linear‑layer outputs across NVIDIA Ampere, Ada, and Hopper GPUs. The new approach achieves 1.17–3.1× faster end‑to‑end performance than existing solutions and halves weight‑memory traffic while maintaining cross‑architecture reproducibility.
By Liam Cooper, Shinnung Jeong, Hyeran Jeon, Jeffrey Young, Hyesoon Kim