TriFleetRCA is an on‑premise pipeline that performs root‑cause analysis for Kubernetes using a single GPU. It gathers evidence at pod, namespace, or cluster scope, deduplicates and ranks it with BM25, filters runbooks through an ingest guard, and returns a root cause with supporting evidence lines. In a live cluster with four injected faults, the system achieved hit rates of 0.85–0.95 across scopes, improved accuracy with deduplication, and demonstrated robust defense against poisoned runbooks.
By Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary
arXiv:2606. 08590v1 Announce Type: cross Abstract: Kubernetes incidents are diagnosed reliably only when a root-cause system's reported gains come from incident evidence rather than scenario-specific shortcuts.
By Anastasiia Kuvshinova, Seungmin Jin
The paper introduces a Kubernetes Dynamic Resource Allocation driver that treats composable CXL memory as a schedulable cluster resource, enabling cross-node shared memory for large language model (LLM) serving. By composing CXL regions on demand, materializing them as DAX devices, and exposing them via a single Container Device Interface name, pods on different nodes can access the same physical memory region. A shared‑memory connector for vLLM/llm‑d uses this region as a KV‑cache tier, eliminating external metadata services and achieving significant reductions in time‑to‑first‑token (TTFT) with minimal additional latency compared to same‑node reuse.
By Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra
arXiv:2606. 03323v2 Announce Type: replace-cross Abstract: The rise of LLM-as-a-Service and other confidential cloud workloads demands cryptographic proof that user data is processed in a trusted, untampered environment.
By Yang Yang, Kevin Wang, Yuanhai Luo, Hang Yin, Jie Cai, Shunfan Zhou, Wenfeng Wang
arXiv:2606. 03323v1 Announce Type: cross Abstract: The rise of LLM-as-a-Service and other confidential cloud workloads demands cryptographic proof that user data is processed in a trusted, untampered environment.
By Yang Yang, Kevin Wang, Yuanhai Luo, Hang Yin, Jie Cai, Shunfan Zhou, Wenfeng Wang
FDE-Bench is a benchmark that tests large language model agents on 136 deployment‑configuration tasks involving Docker, Compose, and Kubernetes, in both greenfield and diagnose‑and‑repair scenarios. Agents submit declarative artifacts that are rebuilt and redeployed in a clean environment, and four binary check layers evaluate build, readiness, behavior, and specification conformance without an LLM judge. The benchmark includes a release gate, detailed check annotations, adversarial strategies, and reports that state‑of‑the‑art models resolve 52.9–75.0 % of tasks, while zero‑intelligence baselines solve none.
By Weihang Ding, Junfei Zhan, Yueting Li, Qirong Guo
arXiv:2604. 16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP).
By Daeyeon Son
arXiv:2607. 21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detection via RFF-approximated Maximum Mean Discrepancy, fairness monitoring with bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a result storage API.
By Lei Yang
arXiv:2606. 30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adversary, validates it without breaking production, and verifies it disrupts the specific attack.
By Chen Frydman, Aviram Zilberman, Rubin Krief, Abed Showgan, Andres Murillo, Sekiya Motoyoshi, Asaf Shabtai, Yuval Elovici, Rami Puzis
arXiv:2608. 07952v1 Announce Type: cross Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging.
By Zhaohui Wang
arXiv:2607. 12723v1 Announce Type: cross Abstract: Filesystem isolation in container ecosystems is often weakened by cross-boundary path misresolution, causing path traversal (PaTra) vulnerabilities.
By Qiyuan Fan, Zhi Li, Junjie Li, XiaoFeng Wang, Bin Yuan, Deqing Zou
Agentic‑Kube is a cooperative multi‑agent reinforcement learning framework for Kubernetes pod placement that splits the multi‑objective scheduling problem into cost minimisation, anti‑affinity fault tolerance, and vector resource balancing, each handled by a dedicated sub‑agent. It uses a bipartite Graph Convolutional Network to model host‑pod dependencies, a two‑stage monotonic QMIX value factorisation network for joint action coherence, and a plurality voting consensus with action feasibility masking. Evaluations on Google Kubernetes Engine and large‑scale clusters show Pareto‑efficient placements, a 53% reduction in anti‑affinity collisions, a 65% spot instance allocation ratio, and sub‑30 ms decision latencies up to 1,000 nodes without container restarts.
By Hamed Hamzeh