arXiv:2603. 16428v2 Announce Type: replace-cross Abstract: Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabilities of most GPUs.
By Ruijia Yang, Zeyi Wen
PipeLive introduces a method for live, in‑place reconfiguration of pipeline parallelism in large language model serving. By redesigning the KV cache layout and extending PageAttention, it enables dynamic resizing of the cache without interrupting inference. The system also uses an incremental KV patching mechanism to keep KV states consistent during reconfiguration, achieving significant reductions in reconfiguration time and improvements in latency metrics.
By Xu Bai, Muhammed Tawfiqul Islam, Chen Wang, Adel N. Toosi
arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.
By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral
The paper proposes a simple technique of chunking workloads into smaller parts that alternate between compute-intensive and memory-bound operations to smooth power and temperature spikes in GPU systems. By doing so, it prevents throttling, leading to faster wall-clock times and lower total energy consumption. Experiments on a DGX Spark show up to 2% performance and energy gains, while similar benefits, though smaller, are observed on multi‑GPU servers.
By Erik Schultheis, Maximilian Kleinegger, Dan Alistarh
arXiv:2606. 19004v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Diffusion Transformers (DiTs) is prohibitively expensive, requiring thousands of high-end GPUs.
By Ruiqi Lai, Dakai An, Wei Gao, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Dmitrii Ustiugov, Wei Wang
arXiv:2606. 00735v1 Announce Type: cross Abstract: In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execution, where the slowest GPU determines layer latency.
By Seokjin Go, Marko Scrbak, Ephrem Wu, Srilatha Manne, Divya Mahajan