arXiv Machine Learning

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

The paper presents an LLM-based predictive scheduling system that forecasts execution time and energy consumption from source code, aiming to improve data center sustainability. By integrating these predictions into a real-time GPU allocation algorithm, the system reduces both energy use and queuing delays. In a collaboration with a data center, the approach achieved a 32% drop in energy consumption and a 30% reduction in waiting time.

arXiv AI
Aug 14

InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers

arXiv:2608. 12915v1 Announce Type: cross Abstract: The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption, and service quality.

By Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
arXiv Machine Learning
Jul 8

Joint Energy Management and Coordinated AIGC Workload Scheduling for Distributed Data Centers: A Diffusion-Aided Reward Shaping Approach

arXiv:2605. 02965v2 Announce Type: replace Abstract: Artificial intelligence-generated content (AIGC) has emerged as a transformative paradigm for automating the creation of diverse and customized content, giving rise to rapidly growing computational workloads in cloud data centers.

By Yang Fu, Peng Qin, Liming Chen, Zihao Zhang, Hao Yu, Yifei Wang
arXiv AI
Sep 4

Artificial Intelligence for Energy Optimization in Data Centers

The paper reviews 194 papers on using artificial intelligence to optimize data center energy use, coding 63 of them. It finds that most control studies validate only in simulation, none consider water withdrawal or embodied carbon, and savings estimates overlap across methods, preventing ranking. The authors propose CLEAR‑DC, a framework that links control and workload demand through elasticity, reports net benefits, and records energy, carbon, water, embodied share, and validation venue.

By Mohammed Basharath Ullah, Summaiya Unnisa Begum, Mohammed Nadeem Ullah
Hugging Face Trending Papers
Jul 29

PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

The rapid growth of AI workloads is turning data centers into large-scale, volatile, yet spatiotemporally flexible grid loads, creating an urgent need for coordinated electricity-computing scheduling. Under stringent grid constraints, schedules from general-purpose large language models (LLMs) are often infeasible, causing line-flow violations and unserved load.

arXiv Machine Learning
5d ago

The Environmental Impacts of Language Model Training Keep Rising Now is the Time to Catch Impacts on the Rebound

The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.

By Cl\'ement Morand (STL), Anne-Laure Ligozat (ENSIIE, LISN, STL), Aur\'elie N\'ev\'eol (STL, LISN)
arXiv AI
Sep 15

Carbon-Aware Routing for Function Calling in Edge-Cloud LLM Systems

The paper presents a carbon‑aware routing framework for function‑calling in large language models that distributes queries across a three‑tier edge‑cloud architecture. A lightweight k‑NN predictor estimates accuracy, delay, and power for each edge tier, and real‑time grid carbon intensity is used to route queries to the lowest‑emission tier that can execute them. Experiments on state‑of‑the‑art benchmarks show the framework matches cloud‑level accuracy while cutting operational carbon emissions by an average of four times.

By Aikaterini Maria Panteleaki, Varatheepan Paramanayakam, Spyros Tragoudas, Iraklis Anagnostopoulos
arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral