The paper presents a method for distributing large language model inference across multiple Intel AI PCs by splitting the model into pipeline shards, each pre‑compiled into an OpenVINO graph. Three key techniques—beam_idx Gather to enable GPU optimizations, speculative decoding on stateful models, and interleaved micro‑batching—allow a two‑node Llama 3.1 8B INT4 pipeline to serve two users at 1.79× the throughput of a single‑node model, while a four‑node deployment can run a 70B model that no single PC can hold. The authors provide code, benchmark logs, and reproduction scripts on GitHub.
By Tate Berenbaum, Muthaiah Venkatachalam
arXiv:2609.37899v1 Announce Type: new
Abstract: Zero-order optimization (ZO) trains without backpropagation, making it relevant to forward-only hardware and non-differentiable loss, but its gradient...
By Francois Chaubard, Mykel J. Kochenderfer, Chris R\'e
arXiv:2605. 25645v2 Announce Type: replace-cross Abstract: We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation.
By Jatin Kishnani, Mayank Goel, Amit Singh, Pulkit Agrawal, Sairanjan Mishra
arXiv:2609.36070v1 Announce Type: cross
Abstract: AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single...
By Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher R\'e
arXiv:2608. 09250v1 Announce Type: new Abstract: Federated learning (FL) must serve devices with varying computational capabilities.
By Bostan Khan, Masoud Daneshtalab
arXiv:2606. 00946v1 Announce Type: cross Abstract: Efficiently serving large language model (LLM) inference tasks is crucial both for user-perceived latency such as time-to-first-token (TTFT) and for GPU utilization.
By Gangmuk Lim, Wanyu Zhao, Brighten Godfrey, Jiaxin Shan, Le Xu, Liguang Xie
arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.
By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao
arXiv:2608. 01563v1 Announce Type: new Abstract: Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries.
By Dzmitry Malyshau
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
By Om Mohite
arXiv:2609.13585v1 Announce Type: cross
Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...
By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica
arXiv:2603. 06741v2 Announce Type: replace-cross Abstract: Training frontier-scale diffusion models often requires substantial computational resources concentrated in tightly-coupled clusters, limiting participation to well-resourced institutions.
By Zhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan Roy
The paper introduces Federation of Experts (FoE), a new architecture that reorganizes the mixture-of-experts (MoE) block in transformer layers into multiple MoE clusters. Each cluster handles a single KV head, and expert parallelism is applied within clusters while a sum operation synchronizes post‑attention residuals across clusters. FoE eliminates all‑to‑all communication on a single GPU and limits it to intra‑node communication in multi‑node setups, leading to significant reductions in inference latency and throughput improvements on LongBench.
By Muhammad Shahir Abdurrahman, Chun Deng, Azalia Mirhoseini, Philip Levis