arXiv AI

dstack-capsule: Pod-Level Remote Attestation for Confidential Workloads on Kubernetes

arXiv:2606. 03323v1 Announce Type: cross Abstract: The rise of LLM-as-a-Service and other confidential cloud workloads demands cryptographic proof that user data is processed in a trusted, untampered environment.

arXiv Machine Learning
Aug 24

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

Thermo-FL is a thermal‑aware federated LoRA fine‑tuning framework that adapts local adapter training and sparse update transmission based on device temperature. It introduces TERRA, a robust aggregation pipeline that uses norm filtering, mask‑aware directional validation, adaptive clipping, and mask‑aware aggregation to defend against Byzantine and communication‑layer adversaries. Experiments on a large‑scale emulator and a Jetson testbed show Thermo-FL improves robustness under adversarial sparse aggregation, stabilizes device temperature, reduces upload size, and preserves utility on GSM8K and BoolQ tasks.

By Shiva Shrestha, Kazi Shaharair Sharif, Zongxing Xie, Jiajing Huang, Anhao Xiang, Honghui Xu
arXiv Machine Learning
Sep 11

Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.

By Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra