Hugging Face Trending Papers

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations leave GPUs idle during pipeline bubbles, wasting computational resources. Asynchronous Pipeline Parallelism eliminates these bubbles, maximizing throughput at the cost of gradient staleness.

Hugging Face Trending Papers
Sep 3

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para‑Pipe is a hierarchical mapping framework that combines intra‑ and inter‑stage operator parallelism within a pipelined architecture to optimize deep‑learning inference on heterogeneous System‑on‑Chip (SoC) platforms. By selectively tuning parallelism levels across pipeline stages, it balances throughput and latency while reducing inter‑processor communication overhead. Evaluations on Amlogic and Black Sesame SoCs show Pareto‑optimal configurations, with throughput‑optimized settings achieving up to 11.0 % higher energy efficiency than purely pipelined approaches and 23.3 % over non‑pipelined parallel execution.

arXiv Machine Learning
Sep 4

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Para-Pipe is a hierarchical mapping framework that integrates intra- and inter-stage operator parallelism within a pipelined architecture for machine‑learning computational graphs on heterogeneous System‑on‑Chip (SoC) platforms. By selectively fine‑tuning parallelism levels across pipeline stages, it navigates the trade‑off between throughput and latency, reducing inter‑processor communication overhead and improving energy efficiency. Evaluation on Amlogic and Black Sesame SoCs shows multiple Pareto‑optimal configurations, with throughput‑optimized setups achieving up to 11.0% better energy efficiency than purely pipelined strategies and 23.3% better than non‑pipelined parallel execution.

By Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra
arXiv AI
Aug 25

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

The paper proposes a theoretical framework for scheduling high‑quality data in large language model training by extending functional scaling laws to account for time‑varying data quality. It identifies two regimes—noise‑limited and signal‑limited—where high‑quality data should be used differently, and introduces a Drop‑Stable‑Rampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.

By Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu
arXiv Machine Learning
Sep 15

Communication-Efficient LLM Adaptation over Decentralized GPU Meshes

The paper introduces a communication‑efficient method for adapting large language models on decentralized GPU meshes. It proposes an asynchronous two‑circuit system that uses fast compressed training with activation masking for pipeline‑parallel transfer and compressed data‑parallel synchronization, while a slower anchor circuit performs occasional unmasked passes. A spectral correction optimizer then denoises the masked gradients using these anchor priors, enabling high compression rates and achieving up to 40× throughput gains over internet‑grade connections while matching dense uncompressed performance.

By Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan, Hadi Mohaghegh Dolatabadi, Chamin P Hewa Koneputugodage, Gil Avraham, Violetta Shevchenko, James Snewin, Karol Pajak, Harry Xi, Alexander Long
arXiv Machine Learning
Jul 21

DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training

arXiv:2604. 26256v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-tailed trajectories indispensable for model performance block the entire training pipeline.

By Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hongyu Zang, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi Gu, Yerui Sun, Yuchen Xie, Xunliang Cai
Hugging Face Trending Papers
Jul 27

Certified Parallel-in-Time Sinkhorn for Dynamic Entropic Optimal Transport

Dynamic applications, including optimal-transport Flow Matching, repeatedly solve related entropic optimal transport problems, yet conventional distributed Sinkhorn processes frames sequentially and synchronizes after every iteration. We present TemporalSinkhorn, a parallel-in-time executor that batches future candidates and their repairs without making output accuracy speculative.