arXiv Machine Learning

Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models

arXiv:2607. 13332v1 Announce Type: new Abstract: Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators, high-speed interconnects, and a single orchestrating entity.

arXiv AI
6d ago

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Scepsy is a serving system designed to efficiently schedule arbitrary multi‑LLM agentic workflows on GPU clusters. It leverages the observation that each LLM’s share of execution time remains relatively stable across requests, profiling LLMs under various parallelism levels to build an Aggregate LLM Pipeline that predicts throughput and latency. Using this predictor, Scepsy searches for optimal GPU allocations—balancing fractional GPU shares, tensor parallelism, and replica counts—and then heuristically places them on the cluster to reduce fragmentation and honor network topology, achieving up to 2.5× higher throughput and 1.0–3.3× lower latency compared to baseline approaches.

By Otto White, Marcel Wagenl\"ander, Britannio Jarrett, Xijin Zhao, Yanda Tao, Pedro Silvestre, Guo Li, Huanzhou Zhu, Llu\'is Vilanova, Peter Pietzuch
arXiv AI
Aug 20

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

The paper presents a method for distributing large language model inference across multiple Intel AI PCs by splitting the model into pipeline shards, each pre‑compiled into an OpenVINO graph. Three key techniques—beam_idx Gather to enable GPU optimizations, speculative decoding on stateful models, and interleaved micro‑batching—allow a two‑node Llama 3.1 8B INT4 pipeline to serve two users at 1.79× the throughput of a single‑node model, while a four‑node deployment can run a 70B model that no single PC can hold. The authors provide code, benchmark logs, and reproduction scripts on GitHub.

By Tate Berenbaum, Muthaiah Venkatachalam
Hugging Face Trending Papers
Sep 17

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By forming loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to traditional sharded DP approaches.

arXiv AI
Sep 15

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...

By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv AI
Sep 18

Accelerating Sharded Data Parallelism at Scale with Federated Learning

The paper proposes two hybrid algorithms, FL+FSDP and FL+HSDP, that combine sharded data parallelism with federated learning-style aggregations to reduce communication overhead in large-scale AI training. By partitioning GPUs into loosely‑coupled federation groups, the methods keep inter‑group traffic minimal while maintaining a bounded global batch size. Experiments on a Llama3.1 8B model trained on 512 A100 GPUs show up to 8.04× faster data processing and 4.48 lower evaluation perplexity compared to conventional sharded DP.

By Gianluca Mittone, Marco Aldinucci
arXiv AI
3d ago

TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training

arXiv:2604.09107v2 Announce Type: replace-cross Abstract: Modern LLM reinforcement learning (RL) workloads require a high-performance weight transfer system to scale training across heterogeneous com...

By Chenhao Ye, Huaizheng Zhang, Mingcong Han, Baoquan Zhong, Xiang Li, Qixiang Chen, Xinyi Zhang, Weidong Zhang, Kaihua Jiang, Wang Zhang, He Sun, Wencong Xiao, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau