arXiv AI

CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift

arXiv:2606. 31470v1 Announce Type: new Abstract: Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency.

arXiv Machine Learning
Sep 4

A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds

A two-stage forecasting system is introduced for predicting CPU workload in private clouds. The model first forecasts customer service requests in Transactions Per Second (TPS) and then estimates future CPU usage from the TPS forecast, both stages using XGBoost within a cascaded architecture. Experiments on real private‑cloud traces show SMAPE below 7% for most applications, with the best case achieving an MAE of 0.7372 and an R² of 0.9185, and stable error accumulation over a 60‑step horizon.

By Ashir Javeed, Anton Borg, H{\aa}kan Grahn, Lars Lundberg, Dhyey Patel, Sogand Shirinbab
arXiv AI
Aug 11

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.

By Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
arXiv AI
Sep 7

Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters

Atlas is a framework that optimizes the deployment of compound AI workflows on heterogeneous clusters by selecting execution plans that satisfy service level objectives (SLOs). It introduces MAP, a Markovian Accuracy Predictor, which estimates configuration accuracy using local conditional accuracy transitions between adjacent workflow stages, avoiding exhaustive end‑to‑end profiling. Atlas formulates plan selection as a mixed‑integer linear program, achieving near‑oracle accuracy while reducing deployment cost by up to 42% and profiling cost by up to 2.6×.

By Milos Gravara, Andrija Stanisic, Stefan Nastic