arXiv Machine Learning

ML-for-ML

arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.

arXiv AI
Aug 19

Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum

The paper introduces a dynamic framework for partitioning neural network layers across a heterogeneous edge‑cloud continuum, adapting to runtime changes in network conditions and device capabilities. It profiles models at startup, measures link quality, and periodically re‑evaluates the partitioning to optimize performance. Experiments on a Raspberry Pi, laptop, and desktop using VGG16, AlexNet, and MobileNetV2 demonstrate energy savings of 27.09–35.82% and latency reductions of 6.34–22.92% over static partitioning.

By Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta
arXiv AI
Sep 15

MANE: A Multi-Path Adaptive Network for Edge Onloading of Deep Neural Networks

MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.

By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris
arXiv Machine Learning
Aug 28

Distributed Training using an Intelligent Network

The paper proposes using the network itself to aid distributed training over a wide area network (WAN). It suggests employing multicast for outbound traffic and in‑line FPGAs for inbound traffic to reduce bottlenecks, extending techniques commonly used in data centers to the WAN. An optimization framework generates synchronization schedules—rotating cliques of compute islands—tailored to the network topology, and demonstrates these ideas on a nine‑city WAN model, showing how schedules adapt to network capabilities.

By Nihar Shah, Ben Blier
arXiv Machine Learning
Aug 31

Ampere: Communication-Efficient and High-Accuracy Split Federated Learning

Ampere is a new split federated learning system that reduces both on‑device computation and device‑server communication while improving accuracy. It trains device and server blocks sequentially with local losses, eliminating gradient transfers, and uses a lightweight auxiliary network to consolidate activations into a single transfer. Experiments on CNNs and Transformers show up to 11.70 pp accuracy gains, 18.6× faster training, 911× less communication, and 14.5× less computation compared to state‑of‑the‑art SFL baselines.

By Zihan Zhang, Leon Wong, Blesson Varghese
arXiv Machine Learning
Sep 7

Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters

The paper introduces REACT, a system that dynamically tunes communication collectives in distributed AI training to mitigate congestion without requiring network infrastructure changes. REACT operates at the application layer, detecting congestion via flow statistics and adjusting the pattern of data exchange—such as selecting different aggregation nodes in an AllReduce tree—while preserving the semantics of the communication. Evaluations on a shared academic GPU cluster show that REACT improves algorithm bandwidth by 13%–38% under congestion, with simulations indicating potential gains up to 75%.

By Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal
arXiv Machine Learning
Sep 2

Contribution-Aware Bandwidth Allocation for Multimodal Split Learning

The paper introduces ModalShare, a bandwidth allocation method for multimodal split learning that assigns each modality a keep‑ratio based on its Shapley contribution score. Unlike existing compression schemes that split the uplink budget proportionally to activation size, ModalShare explicitly optimizes the split across modalities, requiring no extra uplink traffic or client computation. Experiments on CREMA‑D and MVSA datasets show that ModalShare improves accuracy by 12.4–15.4 percentage points over equal keep‑ratios under a 5× compression budget, outperforming three compressors across multiple datasets and budgets.

By Iason Ofeidis, Leandros Tassiulas