arXiv Machine Learning By Caterina Doglioni, Akshat Gupta, Thomas Elliott, Hanzila Hussain, Sanjiban Sengupta

Green BOA: Determining the environmental break-even point for ML-based data compression

Read the original on arXiv Machine Learning →

arXiv:2608. 19994v1 Announce Type: new Abstract: We summarise the outcome of two summer internship projects based at the University of Manchester, focused on the break-even point in terms of environmental sustainability for ML-based data compression algorithms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
4d ago

Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research

The paper highlights that as Large Language Models grow in capability and prevalence, their environmental footprint is increasing, yet the machine learning community lacks standardized carbon accounting practices. An automated review of 5,285 NeurIPS 2025 papers shows almost no reporting of environmental impact. To address this, the authors propose standardized sustainability metrics for training efficiency, heuristics for estimating inference carbon costs, a software tool called carbonbenchmark for tracking emissions, and the SMAJ framework to encourage prioritizing computational efficiency and environmental accountability over marginal accuracy gains.

By Lachlan McGinness, Dan Pagendam, Robert Offner
arXiv AI
Aug 13

Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression

arXiv:2608. 11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression.

By Angelo Nardone, Paolo Ferragina