arXiv Machine Learning

BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training

The paper introduces BASP, a batch‑aware sequence parallelism method that partitions GPUs into disjoint groups based on micro‑batch size to reduce all‑to‑all communication. By localizing communication, BASP improves training efficiency for long‑context LLMs. Experiments on NVIDIA A100 clusters show up to 1.17‑1.31× faster end‑to‑end training on Llama and Qwen models while maintaining the same accuracy and memory usage.

arXiv Machine Learning
Jul 1

HSAP: A Hierarchical Sequence-aware Parallelism for Hybrid-Context Generative Models

arXiv:2606. 30460v2 Announce Type: replace Abstract: In this paper, we aim to combine the advantages of existing sequence parallelism paradigms and overcomes their drawbacks, the most serious of which is the incapability to correctly compute causal attention on the hybrid-context packed sequences, in a stronger sequence parallelism framework.

By Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong lin, Junyu Lu, Jiaxing Zhang, Bingyi Jing
arXiv Machine Learning
Jun 30

HSAP: A Hierachical Sequence-aware Parallelism for Hybrid-Context Generative Models

arXiv:2606. 30460v1 Announce Type: new Abstract: In this paper, we aim to combine the advantages of existing sequence parallelism paradigms and overcomes their drawbacks, the most serious of which is the incapability to correctly compute causal attention on the hybrid-context packed sequences, in a stronger sequence parallelism framework.

By Songxin Zhang, Zejian Xie, Zhuoyang Song, Cong lin, Junyu Lu, Jiaxing Zhang, Bingyi Jing
arXiv Machine Learning
Aug 4

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

arXiv:2605. 08696v4 Announce Type: replace-cross Abstract: Over the last two decades, language modeling has experienced a shift from the use of predominantly recurrent architectures that process tokens sequentially during training and inference to non-recurrent models that process sequence elements in parallel during training, which results in greater training efficiency and stability at the expense of lower inference throughput.

By Benjamin L. Badger
arXiv Machine Learning
Jul 14

Controllably Efficient Language Models

arXiv:2511. 05313v2 Announce Type: replace Abstract: The substantial inference costs of attention in transformers motivated the development of efficient sequence mixers: namely sparse and sliding window attention, convolutions and linear attention.

By Jatin Prakash, Aahlad Puli, Rajesh Ranganath