arXiv Machine Learning

Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering

arXiv:2606. 10896v1 Announce Type: new Abstract: We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass.

arXiv Machine Learning
Aug 18

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

arXiv:2604. 10496v2 Announce Type: replace Abstract: Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling.

By Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa, Valavan Manohararajah, Eric Sather, Sai Qian Zhang
arXiv AI
Aug 18

FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy

arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.

By Qingyao Yang, Runming Yang, He Xiao, Wendong Xu, Junyu Chen, Haobo Liu, Chenchen Ding, Ruihan Hu, Yik-Chung Wu, Ngai Wong
arXiv AI
Jul 7

Panorama: Fast-Track Nearest Neighbors

arXiv:2510. 00566v4 Announce Type: replace-cross Abstract: Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate verification, making it the primary bottleneck in the search process.

By Vansh Ramani, Alexis Schlomer, Akash Nayar, Sayan Ranu, Jignesh M. Patel, Panagiotis Karras
arXiv Machine Learning
Sep 7

Fast Gauss Sums via Flash Attention

The paper introduces a method to compute Gaussian kernel sums—central to many kernel-based machine learning techniques—using flash attention, a highly optimized softmax attention implementation. By applying two small input augmentations, the authors transform the normalized softmax reduction into an unnormalized Gauss sum, eliminating the need for custom GPU code. For feature dimensions greater than eight in fp16, this approach outperforms both compiled PyTorch code and PyKeOps kernels in speed, memory overhead, and accuracy, while maintaining linear memory scaling.

By Nicolaj Rux, Sebastian Neumayer