Purlin: Separating Orchestration from the Datapath of Collectives
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...
arXiv:2607. 28633v1 Announce Type: cross Abstract: Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly.
Scepsy is a serving system designed to efficiently schedule arbitrary multi‑LLM agentic workflows on GPU clusters. It leverages the observation that each LLM’s share of execution time remains relatively stable across requests, profiling LLMs under various parallelism levels to build an Aggregate LLM Pipeline that predicts throughput and latency. Using this predictor, Scepsy searches for optimal GPU allocations—balancing fractional GPU shares, tensor parallelism, and replica counts—and then heuristically places them on the cluster to reduce fragmentation and honor network topology, achieving up to 2.5× higher throughput and 1.0–3.3× lower latency compared to baseline approaches.
arXiv:2504. 08791v3 Announce Type: replace-cross Abstract: On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability.
arXiv:2512. 10236v2 Announce Type: replace-cross Abstract: Modern ML workloads demand distributing training and inference across multiple GPUs.
arXiv:2607. 23264v1 Announce Type: cross Abstract: Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation.