arXiv Machine Learning By Xu Bai, Muhammed Tawfiqul Islam, Chen Wang, Adel N. Toosi

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

Read the original on arXiv Machine Learning →

PipeLive introduces a method for live, in‑place reconfiguration of pipeline parallelism in large language model serving. By redesigning the KV cache layout and extending PageAttention, it enables dynamic resizing of the cache without interrupting inference. The system also uses an incremental KV patching mechanism to keep KV states consistent during reconfiguration, achieving significant reductions in reconfiguration time and improvements in latency metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral