arXiv Machine Learning By Zechen Ma, Zixi Qu, Jinyan Yi, David Lin, Yashar Ganjali

DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training

Read the original on arXiv Machine Learning →

arXiv:2605. 01989v2 Announce Type: replace Abstract: Distributed machine learning (ML) training has become a necessity with the prevalence of billion to trillion-parameter-scale models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 7

ML-for-ML

arXiv:2608. 06046v1 Announce Type: cross Abstract: AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important.

By Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
arXiv AI
Jul 28

Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines

arXiv:2607. 24692v1 Announce Type: cross Abstract: Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predictions.

By Jhonatan Tavori, Gur-Eyal Sela, Ion Stoica, Gil Zussman