Hugging Face Trending Papers

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present TemporalSinkhorn, a parallel-in-time executor that batches future candidates and their repairs without making output accuracy speculative.

arXiv Machine Learning
Sep 4

DrainSinkhorn: Safe Elimination for Batched Entropic Optimal Transport

DrainSinkhorn is a verifier‑gated active‑packing layer that improves batched entropic optimal transport (EOT) by eliminating finished problems from subsequent Sinkhorn updates. It combines candidate‑axis packing, a one‑sided screen, verifier‑gated retirement, and physical compaction, while keeping the EOT objective, per‑instance map, and stopping rule unchanged. The method achieves state‑of‑the‑art execution speedups—up to 4.11× faster on MetroPT‑3 and 3.80× on ImageNet‑32 feature couplings—across multiple backends and tolerance settings. whyItMatters":"The technique delivers significant runtime reductions for heterogeneous batched‑EOT workloads, enabling faster and more efficient optimal transport computations in practical machine‑learning pipelines."

By Xinyang Wen
Hugging Face Trending Papers
Jul 2

DeadPool: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint

State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. Existing fault-tolerance mechanisms either impose non-trivial overhead during failure-free execution or suffer from prolonged recovery latency, particularly under scenarios where a small subset of compute nodes experience permanent failures.

arXiv AI
Jul 28

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference

arXiv:2607. 23264v1 Announce Type: cross Abstract: Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation.

By Jianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song