We Got Claude to Build CUDA Kernels and teach open models!
Related stories
We Got Claude to Fine-Tune an Open Source LLM
Custom Kernels for All from Codex and Claude
Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
arXiv:2607. 24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as matrix multiplication, convolution, and normalization.
Bringing open AI models to the frontier
Why we're building Mistral AI.
How To Build Your Own LLM Runtime From Scratch
If you have ever wanted to actually build an LLM inference runtime yourself — pack your own weights, own every barrier, capture your own CUDA graphs — this is what that journey looks like on an H100. A step-by-step tour of a small runtime called annotated-llm-runtime, and the three bugs that produced most of the annotations.
M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification
arXiv:2603.24595v2 Announce Type: replace-cross Abstract: Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is...
Announcing the OpenAI Learning Accelerator
Welcome GPT OSS, the new open-source model family from OpenAI!
HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization
arXiv:2608.21157v1 Announce Type: cross Abstract: High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolve...
OpenAI Fellows Summer 2018: Final projects
Our first cohort of OpenAI Fellows has concluded, with each Fellow going from a machine learning beginner to core OpenAI contributor in the course of a 6-month apprenticeship.
DICE: Diffusion Large Language Models Excel at Generating CUDA Kernels
arXiv:2602. 11715v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) have emerged as a compelling alternative to autoregressive (AR) LLMs, owing to their capacity for parallel token generation.