arXiv Machine Learning

Mixture-of-Kittens: MoE Megakernel for NVL72s

arXiv AI
Sep 15

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...

By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv AI
Sep 11

KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?

KernelGenBench is a unified benchmark that evaluates large language models and agentic systems for generating efficient Triton kernels across diverse operator sources and hardware platforms. It covers 210 operators from PyTorch ATen, vLLM, and cuBLAS, and tests a 110‑operator subset on six different chips, consuming over 15 billion tokens in evaluation. The study finds that no single method dominates across all sources and platforms, with significant variations in correctness and performance depending on the operator source and hardware, and that agentic approaches require millions of tokens per successful operator.

By Peiyu Zang, Jian Tao, Jialing Zhang, Yichen Yuan, Wentao Zhang, Guang Liu, Yonghua Lin
arXiv Machine Learning
Sep 10

Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems

The paper reports an empirical scalability study of data‑parallel training for Kolmogorov‑Arnold Networks (KANs) on high‑performance computing systems. Using up to eight NVIDIA A100 GPUs across four nodes on the FinisTerrae III supercomputer, the authors evaluate strong and weak scaling, communication overhead, and model‑size scaling, finding a 74.7% parallel efficiency and a 5.97× speedup at eight GPUs. They observe non‑monotonic communication costs driven by All‑Reduce choices and inter‑node latency, and note that while the parameter‑to‑memory ratio improves with larger models, training time scales less favorably, leading to guidelines for GPU topology and model‑size selection.

By Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso
arXiv Machine Learning
Jun 5

CUCo: An Agentic Framework for Compute and Communication Co-design

arXiv:2603. 02376v2 Announce Type: replace-cross Abstract: Computation and communication in distributed LLM training and inference are traditionally optimized in isolation; expert-crafted systems such as DeepEP, FLUX, and TokenWeave show the potential of co-design but require deep systems expertise and hardware-specific tuning; CUCo is an agentic framework that automates compute-communication co-design of CUDA kernels by combining a structured design-space formalization with a correctness-first fast-path agent for reliable baselines and an evolution-driven slow-path agent for high-performance strategies, achieving up to 1.

By Yoga Sri Varshan Varadharajan, Bodun Hu, Saurabh Agarwal, Aditya Akella