arXiv:2607. 24741v1 Announce Type: cross Abstract: Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration.
By Xinyang Wen
DrainSinkhorn is a verifier‑gated active‑packing layer that improves batched entropic optimal transport (EOT) by eliminating finished problems from subsequent Sinkhorn updates. It combines candidate‑axis packing, a one‑sided screen, verifier‑gated retirement, and physical compaction, while keeping the EOT objective, per‑instance map, and stopping rule unchanged. The method achieves state‑of‑the‑art execution speedups—up to 4.11× faster on MetroPT‑3 and 3.80× on ImageNet‑32 feature couplings—across multiple backends and tolerance settings.
whyItMatters":"The technique delivers significant runtime reductions for heterogeneous batched‑EOT workloads, enabling faster and more efficient optimal transport computations in practical machine‑learning pipelines."
By Xinyang Wen
arXiv:2607. 01646v2 Announce Type: replace Abstract: State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack.
By Haotian Xie, Junlin Chen, Mingkai Zheng, Lishan Yang, Zhao Zhang
arXiv:2607. 01646v1 Announce Type: new Abstract: State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack.
By Haotian Xie, Junlin Chen, Mingkai Zheng, Lishan Yang, Zhao Zhang
arXiv:2606. 19004v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Diffusion Transformers (DiTs) is prohibitively expensive, requiring thousands of high-end GPUs.
By Ruiqi Lai, Dakai An, Wei Gao, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Dmitrii Ustiugov, Wei Wang
State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. Existing fault-tolerance mechanisms either impose non-trivial overhead during failure-free execution or suffer from prolonged recovery latency, particularly under scenarios where a small subset of compute nodes experience permanent failures.