Hugging Face Blog

Introducing Storage Regions on the HF Hub

arXiv Machine Learning
4d ago

Characterizing High Bandwidth Flash for LLM Serving

arXiv:2609.39131v1 Announce Type: new Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer,...

By Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
arXiv Machine Learning
Sep 11

Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.

By Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra