arXiv:2607. 18187v1 Announce Type: cross Abstract: Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical.
By Kaiyuan Tang, Maizhe Yang, Chaoli Wang
arXiv:2606. 14353v1 Announce Type: new Abstract: Error-bounded lossy compression is a fundamental technique for managing the rapidly growing volumes of scientific data produced by modern simulations and observational instruments.
By Muhannad Alhumaidi, Guozhong Li, Spiros Skiadopoulos, Panos Kalnis
The paper highlights that as Large Language Models grow in capability and prevalence, their environmental footprint is increasing, yet the machine learning community lacks standardized carbon accounting practices. An automated review of 5,285 NeurIPS 2025 papers shows almost no reporting of environmental impact. To address this, the authors propose standardized sustainability metrics for training efficiency, heuristics for estimating inference carbon costs, a software tool called carbonbenchmark for tracking emissions, and the SMAJ framework to encourage prioritizing computational efficiency and environmental accountability over marginal accuracy gains.
By Lachlan McGinness, Dan Pagendam, Robert Offner
As the capabilities and ubiquity of Large Language Models (LLMs) grow, so does their environmental footprint. Despite calls for responsible AI, the machine learning community lacks standardised practi...
arXiv:2608. 11249v1 Announce Type: cross Abstract: We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - including plain text, source code, and structured formats such as XML - and by recent advances in neural language model-based compression.
By Angelo Nardone, Paolo Ferragina
arXiv:2609.00847v1 Announce Type: cross
Abstract: As machine learning and artificial intelligence find their way into nearly every aspect of climate, weather, and Earth system modeling, it is worth p...
By Filippo Dainelli, Amirpasha Mozaffari, Marina Casta\~no, Aina Gaya i \`Avila, Llu\'is Palma Garcia, Alessio Melli, Oscar Dimdore Miles, Amanda Duarte
arXiv:2606. 05389v1 Announce Type: new Abstract: Lossy compression is essential for massive spatiotemporal data from scientific simulations.
By Liangji Zhu, Sanjay Ranka, Anand Rangarajan
arXiv:2602. 19789v2 Announce Type: replace Abstract: This position paper argues that the machine learning community must move from preaching to practising data frugality for responsible artificial intelligence (AI) development.
By Sophia N. Wilson, Andrew Millard, Gu{\dh}r\'un Fj\'ola Gu{\dh}mundsd\'ottir, Raghavendra Selvan, Sebastian Mair
The paper discusses how any lossless compression algorithm can be transformed into a machine learning method using Normalized Compression Distance or the Minimum Description Length principle, and conversely how any auto‑regressive model can become a lossless compressor via entropy coding. It surveys and formalizes these strategies, introduces a design framework for compression‑based ML, and empirically validates that such methods can match conventional baselines and outperform them on malware detection, achieving accuracy gains up to 0.62 by varying design choices.
By John Hurwitz, Edward Raff, Charles K. Nicholas
The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.
By Cl\'ement Morand (STL), Anne-Laure Ligozat (ENSIIE, LISN, STL), Aur\'elie N\'ev\'eol (STL, LISN)
HyperZip introduces an efficient data compression framework that uses diffusion-based large language models (dLLMs) with Multi-Token Prediction to speed up compression. It addresses the trade‑off between throughput and compression rate by employing a hypernetwork that generates data‑specific updates from a context representation, allowing the dLLM to adapt to target data without costly fine‑tuning. Experiments show HyperZip outperforms state‑of‑the‑art baselines in both compression rate and speed.
By Thai Nguyen, Khang Tran, NhatHai Phan
arXiv:2608. 09998v1 Announce Type: new Abstract: Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks.
By Samar Garrab, Sarra Boughriou, Manel BenSassi