arXiv Machine Learning By Andrew James Amos

From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

Read the original on arXiv Machine Learning →

The paper reports a comprehensive tuning of both a novel sparse self‑organizing map algorithm (SparseBin) and its baseline cuSPARSE implementation. By optimizing four key levers—tile size, tile‑membership clustering, neuron‑axis chunking, and vectorised loads—the authors achieved a 5.6‑10.1× speed‑up per epoch for map sizes ranging from 32×32 to 512×512, and increased the performance margin over the CUDA baseline from ~80× to ~385×. The tuned kernel saturated the L2 bandwidth at 77% of peak, indicating that further performance gains are unlikely without new hardware or fundamentally different approaches.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 26

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

The paper presents a memory‑efficient sparse‑binary self‑organising map (SOM) that scales a MEDLINE atlas to over a million neurons on a single consumer GPU. By re‑ordering the codebook into a feature‑major layout, the authors accelerate the best‑matching‑unit search by 4.5–8.5× without increasing quantisation error, enabling training of a 1,048,576‑neuron SOM in 72 s on a 24 GB GPU. The approach outperforms existing cuSPARSE and CPU‑based SOM implementations, achieving the largest SOM reported to date and demonstrating that resolution limits are computational rather than data‑driven.

By Andrew James Amos
arXiv AI
Aug 26

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

The paper compares two strategies for handling memory limits in large language model (LLM) serving: tensor parallelism, which distributes weights and KV cache across multiple GPUs, and KV compression, which reduces cache size via quantisation and eviction on a single GPU. Using a cost‑normalised simulator calibrated on A100, A40, and H100 hardware, the authors find that across two models (Llama‑2 7B and 70B) and various GPU configurations, compression consistently outperforms tensor parallelism in cost per million tokens, offering 1.20× to 2.00× savings. The study identifies a model‑size threshold (~36B parameters on an 80 GB card) where compression dominates, while tensor parallelism becomes necessary only for larger models where weights alone exceed a single GPU’s capacity.

By Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal, Jay Iyer, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly, Raj Dandekar