Towards Data Science By Anubhab Banerjee

I Built a C++ Backend So My GPU Would Stop Eating Air

Read the original on Towards Data Science →

A comprehensive guide to optimizing LLM inference by eliminating padding overhead with hardware-aware sequence packing. The post I Built a C++ Backend So My GPU Would Stop Eating Air appeared first on Towards Data Science .

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

Towards Data Science
Sep 4

Disaggregation Is a Thousand-GPU Problem

The article discusses the conditions under which separating the prefill phase from the decode phase in large language model inference is beneficial. It explains that only when three specific criteria are met does this split pay off, and it recommends using chunked prefill as the default approach when operating below a certain GPU threshold.

By Mostafa Ibrahim
arXiv Machine Learning
Sep 14

Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

The paper investigates how a single SM utilization metric can misrepresent the true workload of large language model (LLM) inference on Nvidia Hopper GPUs. By profiling vLLM with FlashAttention‑3 and cuBLASLt on an H100 NVL across various phases (cold prefill, warm prefill, and decode) and varying sequence length and batch size, the authors replace the single utilization figure with eight detailed counter‑validated views. These views, tied to specific Nsight Compute counters or formulas, reveal how factors such as fragment fill, occupancy limits, stall signatures, wave quantization, and kernel selection create utilization gaps across four production models and six per‑layer kernel roles.

By Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic, Marco Chiesa