arXiv Machine Learning

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

arXiv:2606. 16384v1 Announce Type: new Abstract: Pretraining language models with extended context windows enhances their ability to leverage rich information during generation.

arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
Hugging Face Trending Papers
Jun 8

End-to-End Context Compression at Scale

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.

arXiv Computation and Language
Sep 25

MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

MILO is a compression framework that reduces the key-value cache memory used in many-shot in-context learning by applying block-wise low-rank compression. It dynamically allocates rank budgets to blocks based on information entropy, preserving important information while aggressively compressing redundant parts. Experiments on Qwen2.5 models show up to a 50% reduction in KV cache memory and a 1.8× throughput improvement with negligible performance loss on classification and reasoning tasks.

By Youpeng Zhao, Tian Tan, Liqian Peng, Jun Wang, Alec Go
arXiv Machine Learning
Aug 4

Structured Recurrent Mixers for Massively Parallelized Sequence Generation

arXiv:2605. 08696v4 Announce Type: replace-cross Abstract: Over the last two decades, language modeling has experienced a shift from the use of predominantly recurrent architectures that process tokens sequentially during training and inference to non-recurrent models that process sequence elements in parallel during training, which results in greater training efficiency and stability at the expense of lower inference throughput.

By Benjamin L. Badger