arXiv:2606. 14707v1 Announce Type: cross Abstract: AI training and deployment consume substantial electricity, but carbon outcomes remain weakly integrated into routine model development decisions.
By Yuxin Chen (University of Helsinki, Finland), Hao Gao (Independent Researcher), Chujie Zou (University of Helsinki, Finland)
arXiv:2605. 23348v2 Announce Type: replace-cross Abstract: AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up.
By Tella Rajashekhar Reddy, Atharva Deshmukh, Liangcheng Yu, Chaojie Zhang, Mike Shepperd, Rohan Gandhi, Anjaly Parayil, Srinivasan Iyengar, Ajay Manchepalli, Debopam Bhattacherjee
The paper presents an LLM-based predictive scheduling system that forecasts execution time and energy consumption from source code, aiming to improve data center sustainability. By integrating these predictions into a real-time GPU allocation algorithm, the system reduces both energy use and queuing delays. In a collaboration with a data center, the approach achieved a 32% drop in energy consumption and a 30% reduction in waiting time.
By Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen
arXiv:2609.05565v1 Announce Type: cross
Abstract: Large language model (LLM) sustainability is increasingly a serving-systems problem, not only a training problem. In production, energy and carbon im...
By Twinkll Sisodia
arXiv:2607. 22806v1 Announce Type: new Abstract: Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power grids.
By Anmol Chaudhary (Department of Electronics,Computer Engineering, NIAMT Ranchi), Rahul Mishra (Department of Electronics,Computer Engineering, NIAMT Ranchi)
The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.
By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos