arXiv Machine Learning By Yifan Niu, Han Xiao, Dongyi Liu, Wei Zhou, Jia Li

Efficient Scaling of LLM Training with Flexible Context Parallelism

Read the original on arXiv Machine Learning →

arXiv:2602. 21788v2 Announce Type: replace-cross Abstract: Scaling long-context capabilities is crucial for Large Language Models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 4

BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training

The paper introduces BASP, a batch‑aware sequence parallelism method that partitions GPUs into disjoint groups based on micro‑batch size to reduce all‑to‑all communication. By localizing communication, BASP improves training efficiency for long‑context LLMs. Experiments on NVIDIA A100 clusters show up to 1.17‑1.31× faster end‑to‑end training on Llama and Qwen models while maintaining the same accuracy and memory usage.

By Bigyan Ghimire, Jon C. Calhoun
arXiv Machine Learning
Jun 16

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

arXiv:2606. 16384v1 Announce Type: new Abstract: Pretraining language models with extended context windows enhances their ability to leverage rich information during generation.

By Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi, Gil Avraham, Violetta Shevchenko, Yan Zuo, Chamin Hewa Koneputugodage, Alexander Long