🤗 Kernels: Major Updates
Related stories
Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
From Zero to GPU: A Guide to Building and Scaling Production-Ready CUDA Kernels
How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
arXiv:2609.21058v1 Announce Type: cross Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a front...
Easily Build and Share ROCm Kernels with Hugging Face
What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
The paper introduces DUALLM, a dual-method pipeline that uses a Large Language Model and a fine‑tuned small language model to classify Linux kernel security patches with high precision. By analyzing commit titles, messages, diffs, and code context, DUALLM achieves 87.4% accuracy and an F1‑score of 0.875, outperforming existing methods. It successfully identified 111 recent patches addressing out‑of‑bounds or use‑after‑free vulnerabilities, with 90 confirmed true positives and proof‑of‑concept exploits demonstrating the validity of the classifications.
KernelGenBench: A Multi-Source and Multi-Chip Benchmark for LLM-based Kernel Generation
arXiv:2607. 27231v1 Announce Type: cross Abstract: Large language models (LLMs) have significantly increased the demand for efficient accelerator kernels, but kernel development remains a highly specialized and labor-intensive task.
Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles
The paper introduces mutation analysis as a metric for evaluating GPU‑kernel benchmark oracles, injecting over ten thousand faults into verified CUDA implementations of 188 KernelBench problems. It shows that the current official checkers miss 16.9% of faults, with precision faults being especially problematic, and demonstrates that optimized test suites can achieve 98% detection with only two inputs per problem. The study also reveals flaws in existing patches and a fuzzing recipe that incorrectly rejects correct kernels 107 times.
M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification
arXiv:2603.24595v2 Announce Type: replace-cross Abstract: Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is...
kAgent: An execution-guided crash resolution agent for the Linux kernel
arXiv:2504. 20412v3 Announce Type: replace-cross Abstract: Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive.
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
KernelGenBench: Can LLMs and Agents Write Efficient Kernels Across Operator Sources and Hardware Platforms?
KernelGenBench is a unified benchmark that evaluates large language models and agentic systems for generating efficient Triton kernels across diverse operator sources and hardware platforms. It covers 210 operators from PyTorch ATen, vLLM, and cuBLAS, and tests a 110‑operator subset on six different chips, consuming over 15 billion tokens in evaluation. The study finds that no single method dominates across all sources and platforms, with significant variations in correctness and performance depending on the operator source and hardware, and that agentic approaches require millions of tokens per successful operator.