arXiv:2607. 24773v1 Announce Type: new Abstract: Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use of physical resources to minimize costs and maximize performance.
By Mehryar Majd, Feng Cheng, Ali Pahlevan
arXiv:2608. 07945v1 Announce Type: cross Abstract: Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge.
By Yifan Wu, Yuhan Li, Zhenhua Wang, Ke Chen, Lidan Shou, Zonghao Chen, Liang Lin, Huan Li, Gang Chen
arXiv:2606. 07565v1 Announce Type: new Abstract: Intelligent scaling of microservices in cloud platforms is crucial for mitigating escalating compute costs while avoiding service disruptions.
By Ahmed Abdulaal, Maruf Aytekin, Thilaga kumaran Srinivasan, Tomer Lancewicki
arXiv:2603.06350v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constrai...
By Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang
arXiv:2606. 13513v1 Announce Type: new Abstract: Driven by conservative over-provisioning to guarantee service reliability, resource utilization in cloud data centers remains at low levels.
By Xiaobin Zhang, Lefei Shen, Mouxiang Chen, Zhuo Li, Hongkai Li, Han Fu, Jianling Sun, Xiaoxue Ren, Chenghao Liu
arXiv:2606. 09787v1 Announce Type: new Abstract: The Cloud-Edge Continuum (CEC) enables latency-critical applications by distributing resources to the far edge, but its extreme volatility makes proactive Zero Touch Management via time-series forecasting essential.
By Abd Elghani Meliani, Arora Sagar, Adlen Ksentini, Raymond Knopp
A two-stage forecasting system is introduced for predicting CPU workload in private clouds. The model first forecasts customer service requests in Transactions Per Second (TPS) and then estimates future CPU usage from the TPS forecast, both stages using XGBoost within a cascaded architecture. Experiments on real private‑cloud traces show SMAPE below 7% for most applications, with the best case achieving an MAE of 0.7372 and an R² of 0.9185, and stable error accumulation over a 60‑step horizon.
By Ashir Javeed, Anton Borg, H{\aa}kan Grahn, Lars Lundberg, Dhyey Patel, Sogand Shirinbab
arXiv:2607. 19974v1 Announce Type: cross Abstract: The rapid proliferation of data-intensive applications, cloud infrastructure, and IoT ecosystems has made proactive resource provisioning critical for maintaining optimal network performance.
By Niraj Gadhe, Kirti Bhardwaj, Moulik Jain, Shubhi Sharma, Vinay Saini
The paper presents an LLM-based predictive scheduling system that forecasts execution time and energy consumption from source code, aiming to improve data center sustainability. By integrating these predictions into a real-time GPU allocation algorithm, the system reduces both energy use and queuing delays. In a collaboration with a data center, the approach achieved a 32% drop in energy consumption and a 30% reduction in waiting time.
By Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen
arXiv:2606. 06687v1 Announce Type: new Abstract: We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers.
By Su Wang, Mung Chiang, H. Vincent Poor
arXiv:2606. 31470v1 Announce Type: new Abstract: Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency.
By Jack Bell, Giacomo Carfi, Gerlando Gramaglia, Andrea Simioni, Daniele Fontani, Vincenzo Lomonaco
arXiv:2608. 14205v1 Announce Type: new Abstract: Load imbalance poses a major bottleneck to the efficiency of expert parallelism in distributed inference of Mixture-of-Experts (MoE) models.
By Pengfei Chen, Yize Wu, Shouxu Kuang, Ke Gao, Ling Li