arXiv AI By Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

Read the original on arXiv AI →

arXiv:2608. 12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 20

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

The paper presents an LLM-based predictive scheduling system that forecasts execution time and energy consumption from source code, aiming to improve data center sustainability. By integrating these predictions into a real-time GPU allocation algorithm, the system reduces both energy use and queuing delays. In a collaboration with a data center, the approach achieved a 32% drop in energy consumption and a 30% reduction in waiting time.

By Hanzhao Wang, Jingxuan Wu, Yumeng Li, Yu Pan, Guanting Chen
arXiv Machine Learning
Jul 28

SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless Clouds

arXiv:2607. 22806v1 Announce Type: new Abstract: Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power grids.

By Anmol Chaudhary (Department of Electronics,Computer Engineering, NIAMT Ranchi), Rahul Mishra (Department of Electronics,Computer Engineering, NIAMT Ranchi)
arXiv AI
Sep 15

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.

By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos