Custom Kernels for All from Codex and Claude
Related stories
🤗 Kernels: Major Updates
Introducing Codex
We Got Claude to Build CUDA Kernels and teach open models!
Easily Build and Share ROCm Kernels with Hugging Face
Codex is now generally available
OpenAI Codex is now generally available with powerful new features for developers: a Slack integration, Codex SDK, and admin tools like usage dashboards and workspace management—making Codex easier to use and manage at scale.
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels
What is Codex?
Learn how Codex helps you go beyond chat by automating tasks, connecting tools, and producing real outputs like docs and dashboards.
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
arXiv:2607. 20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scaling modern ML workloads.
Are you a Codex Original?
We’re collecting real stories of builders, tinkerers, researchers, and creators who are using Codex to do incredible things. If you want to be a part of the next chapter of the Codex Originals program...
Introducing upgrades to Codex
Codex just got faster, more reliable, and better at real-time collaboration and tackling tasks independently anywhere you develop—whether via the terminal, IDE, web, or even your phone.
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
KernelOPT is a multi-agent system that optimizes GPU kernels generated by compilers like PyTorch Inductor by treating compiled models as structured artifacts. It preserves vendor library calls and focuses on Triton sub-kernels, using five profiling-guided LLM agents and a four-gate verification cascade to filter and validate candidates. On 250 KernelBench problems, KernelOPT achieves geometric mean speedups of 1.40×, 1.15×, and 1.07× over torch.compile at different optimization levels.