arXiv AI

Implement Kubernetes Pod-Level Remote Attestation for Confidential Workloads on dstack

arXiv:2606. 03323v2 Announce Type: replace-cross Abstract: The rise of LLM-as-a-Service and other confidential cloud workloads demands cryptographic proof that user data is processed in a trusted, untampered environment.

arXiv Machine Learning
Sep 11

Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.

By Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra