Towards Data Science By Anubhab Banerjee

How To Build Your Own LLM Runtime From Scratch

Read the original on Towards Data Science →

If you have ever wanted to actually build an LLM inference runtime yourself — pack your own weights, own every barrier, capture your own CUDA graphs — this is what that journey looks like on an H100. A step-by-step tour of a small runtime called annotated-llm-runtime, and the three bugs that produced most of the annotations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

arXiv AI
Sep 2

A Composable Evaluation System for Reproducible Omni-Modal Foundation Model Evaluation

The paper introduces OmniEvaluator, a composable evaluation system designed to streamline reproducible testing of omni‑modal foundation models across text, image, video, and audio. It unifies disparate inference engines, prompt conventions, and metric implementations by providing a single interface that supports four inference backends, four evaluation frameworks, and over a thousand benchmarks. Each evaluation run is logged as an artifact for exact reproducibility, with results displayed on a shared dashboard; a federated mode allows GPU inference servers to be shared, and a lightweight verifier ensures stable scoring across engines and prompts without incurring API costs.

By Hodong Lee, Sanghee Park, Dohoon Ryu, Jungwhan Kim, Junyeob Kim, Soyoon Kim, Geewook Kim
arXiv AI
Jun 9

AgentCompile: An LLM-Guided Compiler for Direct CUDA Inference

arXiv:2606. 07665v1 Announce Type: cross Abstract: Transformer inference increasingly depends on specialized compiler and runtime support, but real model graphs still require semantic decisions about which regions are worth specializing and which CUDA implementation families are plausible.

By Xuanzhe Li, Ziyan Weng, Zhiyu Zhu, Junhui Hou