arXiv Machine Learning

Learned Subspace Compression for Communication-Efficient Pipeline Parallelism

arXiv:2606. 05484v1 Announce Type: new Abstract: Pipeline parallelism enables training of large language models that exceed single-device memory, yet inter-stage activation communication becomes the dominant bottleneck when trained on low-bandwidth networks.

arXiv Machine Learning
Jun 16

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

arXiv:2606. 16384v1 Announce Type: new Abstract: Pretraining language models with extended context windows enhances their ability to leverage rich information during generation.

By Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi, Gil Avraham, Violetta Shevchenko, Yan Zuo, Chamin Hewa Koneputugodage, Alexander Long
arXiv AI
2d ago

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

FedLore introduces a communication- and memory-efficient federated learning framework that shares a low-rank optimization basis across clients each round, mitigating subspace fragmentation and enabling exact low-rank aggregation. By refreshing this shared basis across rounds, FedLore allows model updates to exceed the per-round rank budget while maintaining a provable $O(T^{-1/2})$ stationarity bound under standard assumptions. Experiments on vision and language tasks, including federated pre‑training, demonstrate that FedLore outperforms low‑rank adapter baselines and matches or surpasses full‑parameter training while reducing communication and optimizer‑state memory.

By Junkang Liu
arXiv AI
Aug 26

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

The paper introduces the "Compression Trinity," a unified framework that jointly applies sparsity, quantization, and low‑rank approximations to compress large language models. It presents several methods—MKOR, SLoPe, OPTIMA, PATCH, and SLiM—that leverage these three pillars to accelerate training, reduce memory bandwidth, and recover accuracy, achieving significant speedups and accuracy gains over existing techniques. The results demonstrate that combining all three compression strategies is essential for efficient, scalable, high‑performance LLM deployment.

By Mohammad Mozaffari
arXiv Computation and Language
3d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv Computation and Language
Sep 25

MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

MILO is a compression framework that reduces the key-value cache memory used in many-shot in-context learning by applying block-wise low-rank compression. It dynamically allocates rank budgets to blocks based on information entropy, preserving important information while aggressively compressing redundant parts. Experiments on Qwen2.5 models show up to a 50% reduction in KV cache memory and a 1.8× throughput improvement with negligible performance loss on classification and reasoning tasks.

By Youpeng Zhao, Tian Tan, Liqian Peng, Jun Wang, Alec Go
arXiv Machine Learning
Jun 2

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models

arXiv:2606. 00573v1 Announce Type: new Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices.

By Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren, Sai Qian Zhang
arXiv Machine Learning
1d ago

Output-aware Residual Stream Pruning for Large Language Models

The paper proposes a sensitivity‑aware residual‑stream pruning method for large language models that goes beyond minimizing activation reconstruction error. By using a second‑order approximation of output KL divergence, the authors derive a spectral upper bound that selects pruning subspaces based on both activation covariance and output sensitivity, enabling efficient eigendecomposition. Experiments on instruction‑tuned language models show that this approach consistently reduces calibration KL divergence, improves perplexity, and enhances downstream task performance across various compression levels.

By Chayne Thrash, Kevin Chen, Soheil Kolouri