arXiv AI
Jul 29

Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels

arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.

By Joshua Brodsky, Dhravid Kumar, Savini Kashmira, Jayanaka Danatanarayana, Jason Mars, Krisztian Flautner, Lingjia Tang
arXiv Machine Learning
3d ago

Learning Spectrally Optimised Mesh-Free Discretisations

The paper introduces Spectrally Optimised Neural Discretisations (SpeND), a mesh‑free framework that learns discretisation weights from local stencil geometry on unstructured point clouds. By embedding discrete moment conditions into the network architecture, SpeND guarantees polynomial consistency and allows the weights to be optimised for spectral accuracy over a chosen wavenumber band, using an unsupervised Fourier‑mode loss. The resulting operators are PDE‑agnostic, perform well on Poisson, Burgers, and Navier–Stokes equations, and can reduce wall‑clock time by 3–20× compared to existing mesh‑free methods at the same accuracy.

By Lucas Gerken Starepravo, Henry Broadley, Steven Lind, Jack R. C. King
arXiv Machine Learning
Jun 26

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

arXiv:2606. 26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating large language model (LLM) code generation with hardware profiler feedback and pluggable bottleneck detection tools.

By Jiading Gai, Shuai Zhang, Kaj Bostrom, Jin Huang, Vihang Patil, Haoyang Fang, Bernie Wang, Huzefa Rangwala, George Karypis