Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces OrbitTrace, a benchmark of 50 physics‑grounded compute‑availability traces from satellite orbits, and investigates whether specialized interruption‑resilient optimizers are needed when training is interrupted by predictable compute gaps. Experiments on CIFAR‑10/ResNet‑18 and GPT‑2/AdamW show that a strong checkpoint‑and‑resume baseline that preserves full optimizer state and indexes learning‑rate schedules in effective time matches uninterrupted training, rendering most availability‑aware methods unnecessary. Only in a narrow regime—large models with non‑persistable optimizer state and frequent short pauses—does reactive adaptation recover a modest portion of the state‑loss penalty, and even this benefit disappears for eclipse‑scale gaps.
The paper introduces a stability‑aware autoscaling framework for edge serverless workloads that combines an Attention‑Enhanced Double‑Stacked LSTM with Proximal Policy Optimization to address temporal blindness in deep reinforcement learning. By weighting recent historical states non‑uniformly, the method suppresses high‑frequency jitter while preserving demand trends, outperforming single‑layer LSTM, static HPA, and KEDA baselines in latency reduction and stability. Experiments on two Kubernetes clusters with real Azure Functions traces show a ~67% reduction in P90 latency and improved adherence to a 50 ms hard SLO.
arXiv:2607. 21623v1 Announce Type: new Abstract: We present EaaS, a cloud-native reference architecture that operationalizes AI evaluation methods as six stateless Kubernetes microservices: conformal prediction with finite-sample-corrected Adaptive Prediction Sets, calibration assessment, drift detection via RFF-approximated Maximum Mean Discrepancy, fairness monitoring with bootstrap confidence intervals, a DAG-based pipeline orchestrator, and a result storage API.
arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.
arXiv:2609.14762v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy ris...
arXiv:2608. 01725v1 Announce Type: cross Abstract: Modern computing and networking infrastructure emits telemetry continuously, yet operators convert it into decisions with a separate predictor per task, entity, and horizon.