arXiv AI

OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling

arXiv:2509. 23722v2 Announce Type: replace-cross Abstract: Pipeline parallelism is widely used to train large language models (LLMs).

arXiv AI
Jul 7

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

arXiv:2602. 24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distributed serving systems where hundreds of adapters must be hosted concurrently.

By Ferran Agullo, Joan Oliveras, Chen Wang, Alberto Gutierrez-Torre, Olivier Tardieu, Alaa Youssef, Jordi Torres, Josep Ll. Berral
arXiv AI
Jun 9

Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT

arXiv:2601. 20408v2 Announce Type: replace-cross Abstract: Enterprise LLM deployment faces a critical scalability challenge: organizations must optimize models systematically to scale AI initiatives within constrained compute budgets, yet the specialized expertise required for manual optimization remains a niche and scarce skillset.

By Nicholas Santavas, Kareem Eissa, Patrycja Cieplicka, Piotr Florek, Matteo Nulli, Stefan Vasilev, Seyyed Hadi Hashemi, Antonios Gasteratos, Shahram Khadivi
arXiv Machine Learning
Sep 24

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

PipeLive introduces a method for live, in‑place reconfiguration of pipeline parallelism in large language model serving. By redesigning the KV cache layout and extending PageAttention, it enables dynamic resizing of the cache without interrupting inference. The system also uses an incremental KV patching mechanism to keep KV states consistent during reconfiguration, achieving significant reductions in reconfiguration time and improvements in latency metrics.

By Xu Bai, Muhammed Tawfiqul Islam, Chen Wang, Adel N. Toosi
arXiv Machine Learning
Sep 4

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para-Pipe is a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture for machine‑learning computational graphs on heterogeneous System‑on‑Chip (SoC) platforms. By selectively fine‑tuning parallelism levels across pipeline stages, it navigates the trade‑off between throughput and latency, reducing inter‑processor communication overhead and improving energy efficiency. Evaluation on Amlogic and Black Sesame SoCs shows multiple Pareto‑optimal configurations, with throughput‑optimized setups achieving up to 11.0% better energy efficiency than purely pipelined strategies and 23.3% better than non‑pipelined parallel execution.

By Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra