How To Build Your Own LLM Runtime From Scratch
If you have ever wanted to actually build an LLM inference runtime yourself — pack your own weights, own every barrier, capture your own CUDA graphs — this is what that journey looks like on an H100. A step-by-step tour of a small runtime called annotated-llm-runtime, and the three bugs that produced most of the annotations.
Related stories
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
arXiv:2511. 10480v3 Announce Type: replace-cross Abstract: Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed workload execution.
CodegenBench: Can LLMs Write Efficient Code Across Architectures?
arXiv:2606. 04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.
M2K: Making the Model-Kernel Interface Explicit for Reliable CUDA Kernel Verification
arXiv:2603.24595v2 Announce Type: replace-cross Abstract: Large language model (LLM) inference systems rely on CUDA kernels for core GPU computations, yet the interface between models and kernels is...
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
The paper introduces OmniEvaluator, a composable evaluation system designed to streamline reproducible testing of omni‑modal foundation models across text, image, video, and audio. It unifies disparate inference engines, prompt conventions, and metric implementations by providing a single interface that supports four inference backends, four evaluation frameworks, and over a thousand benchmarks. Each evaluation run is logged as an artifact for exact reproducibility, with results displayed on a shared dashboard; a federated mode allows GPU inference servers to be shared, and a lightweight verifier ensures stable scoring across engines and prompts without incurring API costs.
AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference
arXiv:2606. 07665v1 Announce Type: cross Abstract: Transformer inference increasingly depends on specialized compiler and runtime support, but real model graphs still require semantic decisions about which regions are worth specializing and which CUDA implementation families are plausible.
A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation
Building an omni-modal foundation model means evaluating it across text, image, video, and audio. Excellent evaluation toolkits exist for each modality, but their inference engines, prompt conventions...
OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research
arXiv:2607. 16669v1 Announce Type: cross Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible.
DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
DataKernelBench evaluates whether large language models (LLMs) can optimize database queries for GPU execution. The benchmark translates SQL into PyTorch TorchPlan programs and tests LLMs on optimizing core tensor snippets or full queries in CUDA or Triton, using execution-guided repair. On TPC‑H SF10 with an H100 GPU, the best full‑query CUDA configuration outperforms torch.compile by 2.11×, and extending TorchPlan with Dask‑cuDF enables a 2.54× speedup on TPC‑H SF100 across four H100 GPUs.
CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe
arXiv:2604. 01489v2 Announce Type: replace Abstract: High-performance GPU kernels are critical to modern machine learning systems, yet developing them remains a manual, expert-driven process.
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
PerfReasoning is a new benchmark that tests large language models (LLMs) on their ability to reason about hardware performance and generate analytical performance‑model code. The benchmark presents workloads, architectures, and mapping specifications, asking models to compare mappings and predict off‑chip traffic and buffer requirements. While the best closed‑source models achieve over 90% accuracy on reasoning‑based Q&A and the top open‑weight model scores 82.4%, constructing full performance models remains difficult, with most models scoring below 15% and significant variability across runs. Task‑specific reinforcement learning can improve a 4B model’s mapping‑reasoning accuracy by 15.7 points, but feedback‑free self‑revision prompting is not reliably effective.