arXiv AI By Mehryar Majd, Feng Cheng, Ali Pahlevan

Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations

Read the original on arXiv AI →

arXiv:2607. 24773v1 Announce Type: new Abstract: Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs and maximize performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

A Two-Stage Forecasting System for CPU Workload Prediction in Private Clouds

A two-stage forecasting system is introduced for predicting CPU workload in private clouds. The model first forecasts customer service requests in Transactions Per Second (TPS) and then estimates future CPU usage from the TPS forecast, both stages using XGBoost within a cascaded architecture. Experiments on real private‑cloud traces show SMAPE below 7% for most applications, with the best case achieving an MAE of 0.7372 and an R² of 0.9185, and stable error accumulation over a 60‑step horizon.

By Ashir Javeed, Anton Borg, H{\aa}kan Grahn, Lars Lundberg, Dhyey Patel, Sogand Shirinbab
arXiv AI
Aug 11

ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB

arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.

By Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
arXiv Machine Learning
Aug 20

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

The paper presents an LLM-based predictive scheduling system that forecasts execution time and energy consumption from source code, aiming to improve data center sustainability. By integrating these predictions into a real-time GPU allocation algorithm, the system reduces both energy use and queuing delays. In a collaboration with a data center, the approach achieved a 32% drop in energy consumption and a 30% reduction in waiting time.

By Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen