arXiv Machine Learning By Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo, Alexander Long

Subspace Networks: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Read the original on arXiv Machine Learning →

arXiv:2506. 01260v3 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

Mixtures of Subspaces for Bandwidth Efficient Context Parallel Training

arXiv:2606. 16384v1 Announce Type: new Abstract: Pretraining language models with extended context windows enhances their ability to leverage rich information during generation.

By Sameera Ramasinghe, Ajanthan Thalaiyasingam, Hadi Mohaghegh Dolatabadi, Gil Avraham, Violetta Shevchenko, Yan Zuo, Chamin Hewa Koneputugodage, Alexander Long
arXiv AI
Sep 2

A Mathematical Theory of Reusable Neural Bases for Network Compression

The paper introduces the Linear Reusable Neural Bases Architecture (LRNBA), a framework that represents each network block as a linear combination of shared neural bases to improve parameter efficiency and reduce memory cost. Inspired by recurrent neural network designs, LRNBA enables the construction of wider and deeper networks within the same parameter budget. Experiments show that models using LRNBA converge as fast or faster than classical architectures, achieve lower loss, and maintain stable training dynamics.

By Binshuai Wang