arXiv AI By Sirui Chen, Jingji Chen, Siqi Zhu, Ziheng Jiang, Yanghua Peng, Xuehai Qian

Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality

Read the original on arXiv AI →

arXiv:2512. 20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited parallelism or incur high communication costs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.