Case Study: Millisecond Latency using Hugging Face Infinity and modern CPUs
Related stories
Intel and Hugging Face Partner to Democratize Machine Learning Hardware Acceleration
Hugging Face Models on Foundry Managed Compute
Accelerating over 130,000 Hugging Face models with ONNX Runtime
Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
arXiv:2608. 10103v1 Announce Type: cross Abstract: High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
arXiv:2607. 06202v1 Announce Type: cross Abstract: The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth.
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
The paper introduces Decode‑Latency Feedback Prefill (DLFP), a model‑free controller that adjusts prefill chunk sizes during concurrent autoregressive inference to reduce interference between new and ongoing requests. Implemented in vLLM, DLFP achieves significant reductions in P99 inter‑token latency on Qwen3‑0.6B while maintaining output correctness and SLO compliance, though it fails to generalize to larger models or multi‑GPU setups. The study highlights the limits of this approach and suggests the need for a completion‑timed controller for broader applicability.
CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems
arXiv:2607. 07862v1 Announce Type: cross Abstract: The evolution of compute infrastructure has transformed multi-GPU systems into tightly integrated shared-memory structures.
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
arXiv:2607. 02640v1 Announce Type: cross Abstract: Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingests streaming audio and must respond by a recurring wall-clock deadline, while its KV cache grows monotonically and stays pinned for the whole conversation.
X-Stage: Modeling Post-Issue Backpressure in GPU Communication--Computation Fusion
arXiv:2607.23264v2 Announce Type: replace-cross Abstract: Fine-grained, device-initiated communication allows fused GPU kernels to issue remote stores directly from their compute pipelines, a pattern...
Fetch Cuts ML Processing Latency by 50% Using Amazon SageMaker & Hugging Face
Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management
The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.