Efficient Scaling of LLM Training with Flexible Context Parallelism
arXiv:2602. 21788v2 Announce Type: replace-cross Abstract: Scaling long-context capabilities is crucial for Large Language Models (LLMs).
arXiv:2606. 08476v1 Announce Type: cross Abstract: Context parallelism (CP) is essential for training large-scale, long-context language models, as it partitions sequences to reduce memory overhead.
arXiv:2602. 21788v2 Announce Type: replace-cross Abstract: Scaling long-context capabilities is crucial for Large Language Models (LLMs).
arXiv:2511. 10480v3 Announce Type: replace-cross Abstract: Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed workload execution.
arXiv:2606. 11169v1 Announce Type: cross Abstract: Large-scale model training increasingly relies on composing multiple parallelism strategies, such as data, pipeline, and expert parallelism, together with memory-saving optimizations like ZeRO.
The paper introduces Faster Flash Decoding (FFD), a hardware‑algorithm co‑design that fuses the selector and computation into a single kernel to eliminate memory‑bandwidth bottlenecks in long‑context decoding. By replacing metadata indices with low‑bit quantized, content‑aware scanning and employing a top‑delta strategy for dynamic block filtering, FFD achieves up to 11.6× kernel‑level speedup and scales to 256K context length while preserving model accuracy. The approach is training‑free, plug‑and‑play, and demonstrates significant throughput gains on benchmarks such as RULER and LongBench.
The paper introduces BASP, a batch‑aware sequence parallelism method that partitions GPUs into disjoint groups based on micro‑batch size to reduce all‑to‑all communication. By localizing communication, BASP improves training efficiency for long‑context LLMs. Experiments on NVIDIA A100 clusters show up to 1.17‑1.31× faster end‑to‑end training on Llama and Qwen models while maintaining the same accuracy and memory usage.
DeepSeek‑V4.1‑Flash is a multimodal Mixture‑of‑Experts model with 552 B backbone parameters that supports contexts of up to one million tokens. It uses a Causal Encoder‑Decoder architecture that activates 16 B parameters per token during decoding but only 8 B during prefill, improving cost efficiency for agentic workloads. The model introduces cross‑layer KV cache reuse via Compressed Sparse Attention 2 and FP4 KV caching, reducing its global KV cache footprint to 890 bytes per token and its persistent footprint to roughly one‑eighth of the previous version, while still delivering superior performance after pretraining on a 45 T‑token multimodal corpus.
arXiv:2512. 22219v2 Announce Type: replace-cross Abstract: We introduce Mirage Persistent Kernel (MPK), the first compiler and runtime system that automatically transforms multi-GPU model inference into a single high-performance mega-kernel.
arXiv:2608. 07974v1 Announce Type: new Abstract: Large language model (LLM) fine-tuning at the edge adapts the model to scenario-specific data while preserving privacy.
arXiv:2606. 30460v2 Announce Type: replace Abstract: In this paper, we aim to combine the advantages of existing sequence parallelism paradigms and overcomes their drawbacks, the most serious of which is the incapability to correctly compute causal attention on the hybrid-context packed sequences, in a stronger sequence parallelism framework.
arXiv:2602. 21196v2 Announce Type: replace Abstract: Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism.
arXiv:2608. 19758v1 Announce Type: new Abstract: Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase.
arXiv:2607. 08057v1 Announce Type: cross Abstract: Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly.