arXiv:2402. 09589v2 Announce Type: replace-cross Abstract: We present MLCC, a novel technique to augment today's congestion control algorithms to accelerate DNN training jobs in shared GPU clusters in a fully distributed manner.
By Anton A. Zabreyko, Sanjoli Narang, Sudarsanan Rajasekaran, Manya Ghobadi
The paper introduces a dynamic framework for partitioning neural network layers across a heterogeneous edge‑cloud continuum, adapting to runtime changes in network conditions and device capabilities. It profiles models at startup, measures link quality, and periodically re‑evaluates the partitioning to optimize performance. Experiments on a Raspberry Pi, laptop, and desktop using VGG16, AlexNet, and MobileNetV2 demonstrate energy savings of 27.09–35.82% and latency reductions of 6.34–22.92% over static partitioning.
By Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta
MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.
By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris
arXiv:2605. 01989v2 Announce Type: replace Abstract: Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillion-parameter-scale models.
By Zechen Ma, Zixi Qu, Jinyan Yi, David Lin, Yashar Ganjali
arXiv:2507. 09029v5 Announce Type: replace-cross Abstract: Pre-training large neural networks at scale imposes heavy memory demands on accelerators and often requires costly communication.
By Vaibhav Singh, Zafir Khalid, Pietro Cagnasso, Edouard Oyallon, Eugene Belilovsky
arXiv:2512.04705v3 Announce Type: replace-cross
Abstract: The deployment of Early-Exiting Neural Networks (EENNs) on edge accelerators requires optimizing not only the network architecture but also i...
By Alaa Zniber, Arne Symons, Ouassim Karrakchou, Marian Verhelst, Mounir Ghogho