← Back to all news
arXiv Machine Learning September 30, 2026 By Peng Xu, Nihar Koganti, Volodymyr Kindratenko, Xiaohui Chen

GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 28

Low-Rank Dependence Decomposition via Accelerated Symmetric Non-negative Matrix Factorization

arXiv:2607. 24518v1 Announce Type: new Abstract: Symmetric non-negative matrix factorization (SymNMF) recovers latent group structure from a dependence matrix, but its dense, quadratic-memory objective has confined prior work to moderate sizes.

By Lavinia Ghita, Dhruv Desai, Jake Goldberg, Roman Yokunda Enzmann
benchmarks
More like this →
arXiv Machine Learning
Jun 10

Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering

arXiv:2606. 10896v1 Announce Type: new Abstract: We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass.

By Gal Bloch, Ariel Gera, Matan Orbach, Ohad Eytan, Assaf Toledo
More like this →
arXiv AI
Aug 11

Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality

arXiv:2512. 20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited parallelism or incur high communication costs.

By Sirui Chen, Jingji Chen, Siqi Zhu, Ziheng Jiang, Yanghua Peng, Xuehai Qian
llms
More like this →
arXiv AI
Jul 13

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

arXiv:2607. 08786v1 Announce Type: cross Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge.

By Tao Lu, Haoyu Wang, Zonghui Wang, Keshen Xiang, Jiaheng Zhang, Wenzhi Chen
llmsefficiency
More like this →
arXiv Machine Learning
Jun 15

Concatenated Matrix SVD: Compression Bounds, Incremental Approximation, and Error-Constrained Clustering

arXiv:2601. 11626v2 Announce Type: replace-cross Abstract: Large collections of matrices arise throughout modern machine learning, signal processing, and scientific computing, where they are commonly compressed by concatenation followed by truncated singular value decomposition (SVD).

By Maksym Shamrai
More like this →
arXiv Machine Learning
Sep 10

GPU-Enabled Large-Scale Optimization Using Randomized Linear Algebra

arXiv:2609. 08136v1 Announce Type: new Abstract: This paper introduces rlaopt, a PyTorch-based package for large-scale optimization and scientific computing using randomized numerical linear algebra (RandNLA).

By Pratik Rathore, Zachary Frangella, Parth Nobel, Xuning Hu, Madeleine Udell
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea