arXiv Machine Learning

Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing

arXiv Machine Learning
Sep 22

When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules

The paper introduces OrbitTrace, a benchmark of 50 physics‑grounded compute‑availability traces from satellite orbits, and investigates whether specialized interruption‑resilient optimizers are needed when training is interrupted by predictable compute gaps. Experiments on CIFAR‑10/ResNet‑18 and GPT‑2/AdamW show that a strong checkpoint‑and‑resume baseline that preserves full optimizer state and indexes learning‑rate schedules in effective time matches uninterrupted training, rendering most availability‑aware methods unnecessary. Only in a narrow regime—large models with non‑persistable optimizer state and frequent short pauses—does reactive adaptation recover a modest portion of the state‑loss penalty, and even this benefit disappears for eclipse‑scale gaps.

By Subhadip Mitra
arXiv Machine Learning
Sep 24

Learning to Remember: Attentive Reinforcement Learning for Edge Serverless Autoscaling

The paper introduces a stability‑aware autoscaling framework for edge serverless workloads that combines an Attention‑Enhanced Double‑Stacked LSTM with Proximal Policy Optimization to address temporal blindness in deep reinforcement learning. By weighting recent historical states non‑uniformly, the method suppresses high‑frequency jitter while preserving demand trends, outperforming single‑layer LSTM, static HPA, and KEDA baselines in latency reduction and stability. Experiments on two Kubernetes clusters with real Azure Functions traces show a ~67% reduction in P90 latency and improved adherence to a 50 ms hard SLO.

By Faraz Shaikh, Gianluca Reali, Mauro Femminella
arXiv Machine Learning
Jul 27

Cloud-Native Evaluation-as-a-Service: A Microservices Architecture for Scalable AI Monitoring with Conformal Guarantees

arXiv:2607. 21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detection via RFF-approximated Maximum Mean Discrepancy, fairness monitoring with bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a result storage API.

By Lei Yang
arXiv Machine Learning
Aug 4

Real-Time Detection and Repair of LLM Agent Failures

arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.

By Sunny Dubey
arXiv Machine Learning
Sep 22

TriFleetRCA: On-Premise LLM Root Cause Analysis for Kubernetes

TriFleetRCA is an on‑premise pipeline that performs root‑cause analysis for Kubernetes using a single GPU. It gathers evidence at pod, namespace, or cluster scope, deduplicates and ranks it with BM25, filters runbooks through an ingest guard, and returns a root cause with supporting evidence lines. In a live cluster with four injected faults, the system achieved hit rates of 0.85–0.95 across scopes, improved accuracy with deduplication, and demonstrated robust defense against poisoned runbooks.

By Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary
arXiv Machine Learning
Jun 11

NetBurst: Event-Centric Forecasting of Bursty, Intermittent Time Series

arXiv:2510. 22397v2 Announce Type: replace-cross Abstract: Network operators monitor their infrastructure by collecting telemetry data such as packet counts, byte rates, or flow volumes, yet answering the questions that effective operations demand -- forecasting future load, diagnosing and characterizing anomalies, and searching for and retrieving historical precedents -- requires more than raw measurements.

By Satyandra Guthula, Jaber Daneshamooz, Charles Fleming, Kesheng Wu, Walter Willinger, Arpit Gupta
arXiv AI
Aug 26

Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap

The paper introduces an elastic key‑value (KV) cache for large language model (LLM) serving that dynamically reclaims a pre‑allocated reserve during decode‑heavy phases and restores it before prefill, using a userspace CUDA virtual‑memory trick that requires no driver changes. The authors implement this mechanism, test it under realistic workloads, and find that it offers only marginal benefits—about a 1 % difference in time‑to‑first‑token for large prefill chunks—and that simpler strategies such as lowering the maximum batch size can achieve similar results. The study also notes that the reserve’s impact diminishes with higher tensor‑parallelism levels. whyItMatters":"The work demonstrates that a dynamic KV cache reclamation strategy can be implemented without driver patches and that its practical benefits are limited, guiding future LLM serving optimizations toward simpler approaches."

By Sathishkumar Sivashanmugam
arXiv AI
Sep 24

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Crossflow introduces an elastic boundary for prefilling and decoding in large language model serving, allowing decode nodes to publish short‑lived leases that limit prefilling resources and output projections. By adapting to dynamic phase demand, Crossflow improves token throughput by 16.2‑17.4% on average and up to 43.4% under high load, while consistently reducing mean time‑to‑first‑token. The approach eliminates the inefficiencies of static partitioning, which can leave 17% of cluster capacity idle or cause queueing and lost throughput.

By Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang