arXiv:2608. 12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality.
By Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
arXiv:2606. 14707v1 Announce Type: cross Abstract: AI training and deployment consume substantial electricity, but carbon outcomes remain weakly integrated into routine model development decisions.
By Yuxin Chen (University of Helsinki, Finland), Hao Gao (Independent Researcher), Chujie Zou (University of Helsinki, Finland)
The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.
By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos
arXiv:2605. 23348v2 Announce Type: replace-cross Abstract: AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up.
By Tella Rajashekhar Reddy, Atharva Deshmukh, Liangcheng Yu, Chaojie Zhang, Mike Shepperd, Rohan Gandhi, Anjaly Parayil, Srinivasan Iyengar, Ajay Manchepalli, Debopam Bhattacherjee
arXiv:2609.33965v2 Announce Type: replace-cross
Abstract: We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between i...
By Joshua Horswill, Ross Hunter, Matt Clifford, James Hall
arXiv:2606. 30919v1 Announce Type: cross Abstract: Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models at the edge to stronger models in the cloud.
By Wei Geng, Nitinder Mohan, J\"org Ott