arXiv:2608.28910v1 Announce Type: new
Abstract: Coded computing has emerged as a powerful paradigm for mitigating the impact of straggling workers in distributed computing systems. However, existing...
By Parsa Moradi, Behrooz Tahmasebi, Mohammad Ali Maddah-Ali
arXiv:2606. 05484v1 Announce Type: new Abstract: Pipeline parallelism enables training of large language models that exceed single-device memory, yet inter-stage activation communication becomes the dominant bottleneck when trained on low-bandwidth networks.
By Paul Janson, Edouard Oyallon, Eugene Belilovsky
arXiv:2508. 06588v3 Announce Type: replace-cross Abstract: Vector Quantization (VQ) has recently emerged as a promising approach for learning compressed and discrete representations for graph-structured data.
By Zian Zhai, Fan Li, Xingyu Tan, Xiaoyang Wang, Wenjie Zhang
arXiv:2607. 11883v1 Announce Type: new Abstract: Compression is fundamental to intelligence.
By Shikai Qiu, Marc Finzi, Yujia Zheng, Kun Zhang, Andrew Gordon Wilson
arXiv:2608.29867v1 Announce Type: new
Abstract: Autoencoders are widely used for nonlinear dimensionality reduction and manifold learning. While most common implementations rely on both nonlinear enc...
By Louen Pottier, Louis Lesueur, Anders Thorin
arXiv:2602.15239v3 Announce Type: replace
Abstract: Transformers have achieved remarkable success across domains, motivating the rise of Graph Transformers (GTs) as attention-based architectures for...
By Javier Porras-Valenzuela, Zhiyang Wang, Teresa Shang, Yusu Wang, Alejandro Ribeiro
arXiv:2506. 01260v3 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks.
By Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo, Alexander Long
Compression is fundamental to intelligence. A model that can represent its training data as a short code has discovered regularities that enable generalization.
GraphK introduces an encoder‑sampler‑decoder framework that generates variable‑size graphs efficiently. It learns permutation‑invariant latent representations and samples new node embeddings via maximum likelihood, enabling both upscaling and downscaling of graph size. Edge construction uses KDTree‑based top‑k neighbor search in latent space, reducing computational cost while capturing graph properties.
By Resul Tugay, Eren Olu\u{g}, Elif Ak, Sule Gunduz Oguducu
The paper introduces COSA-GS, a new compression method for 3D Gaussian Splatting that avoids spatial aggregation by using anchor-wise causal factorization. It builds a compact learnable anchor latent from geometry context and fuses it with the geometry context to create an anchor context for attribute coding, employing only linear transformations and activations. The method is trained with rate–distortion optimization, adaptive Gaussian pruning, and quantization-aware training to ensure bit‑exact entropy decoding across platforms, achieving state‑of‑the‑art compression performance with fast, consistent cross‑platform decoding.
By Pengpeng Yu, Yueru Chen, Fei Song, Tai Qin, Qi Zhang, Jing Wang, Yulan Guo
arXiv:2506. 01260v2 Announce Type: replace Abstract: Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks.
By Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo, Alexander Long
The paper introduces LA-ReduNet, a lightweight adaptive version of the ReduNet neural network that uses hyperspherical manifold learning and adaptive step sizes to reduce the number of layers needed for the maximal coding rate reduction (MCR$^2$) objective to stabilize. By refining the layer‑wise update rule, LA-ReduNet achieves comparable classification accuracy while requiring far fewer layers and significantly less parameter storage—about 1/29 of the unfolded ReduNet module under the tested settings.
By Zhenglin Huang, Qifa Yan, Bin Dai, Xiaohu Tang