arXiv Machine Learning By Miguel Braga, J\'ulio Pinto, Rahma Nouaji, Olivier Michaud, Bettina Kemme, Oana Balmau, Cl\'audia Brito, Ricardo Macedo

A principled approach for energy-efficient training via phase-aware GPU frequency tuning

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

arXiv Machine Learning
Sep 14

One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training

The paper proposes a simple technique of chunking workloads into smaller parts that alternate between compute-intensive and memory-bound operations to smooth power and temperature spikes in GPU systems. By doing so, it prevents throttling, leading to faster wall-clock times and lower total energy consumption. Experiments on a DGX Spark show up to 2% performance and energy gains, while similar benefits, though smaller, are observed on multi‑GPU servers.

By Erik Schultheis, Maximilian Kleinegger, Dan Alistarh
arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral