arXiv AI

The Environmental Cost of LLMs in AIED: Reporting and Practices

arXiv:2606. 11215v1 Announce Type: cross Abstract: Large Language Model (LLM) usage in recent years has become increasingly widespread in the Artificial Intelligence in Education (AIED) community.

arXiv AI
2d ago

Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research

The paper highlights that as Large Language Models grow in capability and prevalence, their environmental footprint is increasing, yet the machine learning community lacks standardized carbon accounting practices. An automated review of 5,285 NeurIPS 2025 papers shows almost no reporting of environmental impact. To address this, the authors propose standardized sustainability metrics for training efficiency, heuristics for estimating inference carbon costs, a software tool called carbonbenchmark for tracking emissions, and the SMAJ framework to encourage prioritizing computational efficiency and environmental accountability over marginal accuracy gains.

By Lachlan McGinness, Dan Pagendam, Robert Offner
arXiv Machine Learning
Sep 18

The Environmental Impacts of Language Model Training Keep Rising Now is the Time to Catch Impacts on the Rebound

The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.

By Cl\'ement Morand (STL), Anne-Laure Ligozat (ENSIIE, LISN, STL), Aur\'elie N\'ev\'eol (STL, LISN)
arXiv AI
Aug 19

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

The paper explores energy-aware knowledge distillation for large language models (LLMs) used in software engineering tasks such as clone detection, vulnerability prediction, and code summarization. It shows that the commonly used FLOPs metric does not reliably reflect actual energy consumption, and that using energy-surrogate models during distillation can reduce inference energy by up to 90% and memory usage by 86% with only modest accuracy loss. The study demonstrates that guiding distillation with direct energy estimates improves the sustainability and deployability of LLMs on consumer hardware.

By Enrique Barba Roque, Lu\'is Cruz, Annibale Panichella
arXiv Computation and Language
Sep 10

$\Phi$-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

arXiv:2609.10226v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in...

By Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan, Yinmin Zhang, Qi Han, Yiming Xu, Feiyuan Zhang, Xiaomeng Chu, Guoliang You, Wuyang Zhang, Daxin Jiang, Yanyong Zhang
arXiv Machine Learning
Jul 8

Life Cycle Assessment of Pre-training the Lucie 7B Open-Source Large Language Model on the Jean Zay Supercomputer

arXiv:2607. 05408v1 Announce Type: cross Abstract: The environmental impact of training large language models (LLMs) is increasingly scrutinised, yet most published estimates focus on operational energy and disclose little about manufacturing (embodied) emissions, water consumption, or the underlying highperformance computing (HPC) infrastructure.

By Marc L\'eobet, Pierre-Fran\c{c}ois Lavall\'ee, Jean-Pierre Lorr\'e