Introducing Storage Regions on the HF Hub
Related stories
Designing the hf CLI as an agent-optimized way to work with the Hub
Announcing Evaluation on the Hub
Introducing Storage Buckets on the Hugging Face Hub
Welcome RL Environments to the hub
From Chunks to Blocks: Accelerating Uploads and Downloads on the Hub
Run AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilot
Characterizing High Bandwidth Flash for LLM Serving
arXiv:2609.39131v1 Announce Type: new Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer,...
How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes
Architecting cost-effective infrastructure by navigating the latency and storage trade-offs of HNSW, SPANN, and DiskANN The post How to Optimize Vector Search When RAM Gets Too Expensive: On-Disk vs. In-Memory ANN Indexes appeared first on Towards Data Science .
Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving
The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.
RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving
arXiv:2609.07008v1 Announce Type: new Abstract: Multi-head latent attention (MLA) exposes many logical query heads through one packed latent KV stream. This representation is memory efficient, but it...