arXiv AI

Zephon: Elastic Determinism for Online, Stateful Foundation Model Data Loading Pipelines

Zephon is a data loader designed for foundation model training that ensures deterministic ordering of training data batches even when GPU resources change, checkpoints are resumed, or different processing backends are used. It handles online, stateful pipelines—where tokenization, packing, and mixing of samples create complex n‑to‑m transformations—by partitioning the data stream into topology‑independent lanes, serializing ordering decisions, and checkpointing only bounded in‑flight state. Experiments on text and vision‑language tasks show Zephon delivers competitive throughput while offering guarantees that existing loaders lack for such pipelines.

arXiv AI
Jun 29

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers

arXiv:2601. 16956v1 Announce Type: cross Abstract: The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has necessitated training across thousands of GPUs using complex hybrid parallelism strategies (e.

By Avinash Maurya, M. Mustafa Rafique, Franck Cappello, Bogdan Nicolae
arXiv AI
6d ago

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training

arXiv:2604.09107v2 Announce Type: replace-cross Abstract: Modern LLM reinforcement learning (RL) workloads require a high-performance weight transfer system to scale training across heterogeneous com...

By Chenhao Ye, Huaizheng Zhang, Mingcong Han, Baoquan Zhong, Xiang Li, Qixiang Chen, Xinyi Zhang, Weidong Zhang, Kaihua Jiang, Wang Zhang, He Sun, Wencong Xiao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
arXiv AI
Sep 25

TIDE: Temporal Incremental Draft Engine for Self-Improving LLM Inference

TIDE (Temporal Incremental Draft Engine) is a serving‑engine‑native framework that integrates online draft adaptation into high‑performance LLM inference. By reusing intermediate hidden states from the target model as training signals, TIDE avoids extra target model computation and serving‑time overhead, activating speculation and draft training only when beneficial. On heterogeneous GPU clusters, TIDE achieves up to 1.66× higher throughput than no‑speculation baselines, reduces training time by up to 3.02×, cuts storage needs by 24×, and improves system throughput by up to 1.22×.

By Jiyoung Park, Hankyu Jang, Changseok Song, Wookeun Jung
arXiv Machine Learning
Sep 24

PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving

PipeLive introduces a method for live, in‑place reconfiguration of pipeline parallelism in large language model serving. By redesigning the KV cache layout and extending PageAttention, it enables dynamic resizing of the cache without interrupting inference. The system also uses an incremental KV patching mechanism to keep KV states consistent during reconfiguration, achieving significant reductions in reconfiguration time and improvements in latency metrics.

By Xu Bai, Muhammed Tawfiqul Islam, Chen Wang, Adel N. Toosi
arXiv AI
Aug 20

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

The paper presents a method for distributing large language model inference across multiple Intel AI PCs by splitting the model into pipeline shards, each pre‑compiled into an OpenVINO graph. Three key techniques—beam_idx Gather to enable GPU optimizations, speculative decoding on stateful models, and interleaved micro‑batching—allow a two‑node Llama 3.1 8B INT4 pipeline to serve two users at 1.79× the throughput of a single‑node model, while a four‑node deployment can run a 70B model that no single PC can hold. The authors provide code, benchmark logs, and reproduction scripts on GitHub.

By Tate Berenbaum, Muthaiah Venkatachalam