arXiv Machine Learning By Andrew James Amos

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

Read the original on arXiv Machine Learning →

The paper presents a memory‑efficient sparse‑binary self‑organising map (SOM) that scales a MEDLINE atlas to over a million neurons on a single consumer GPU. By re‑ordering the codebook into a feature‑major layout, the authors accelerate the best‑matching‑unit search by 4.5–8.5× without increasing quantisation error, enabling training of a 1,048,576‑neuron SOM in 72 s on a 24 GB GPU. The approach outperforms existing cuSPARSE and CPU‑based SOM implementations, achieving the largest SOM reported to date and demonstrating that resolution limits are computational rather than data‑driven.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 7

From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

The paper reports a comprehensive tuning of both a novel sparse self‑organizing map algorithm (SparseBin) and its baseline cuSPARSE implementation. By optimizing four key levers—tile size, tile‑membership clustering, neuron‑axis chunking, and vectorised loads—the authors achieved a 5.6‑10.1× speed‑up per epoch for map sizes ranging from 32×32 to 512×512, and increased the performance margin over the CUDA baseline from ~80× to ~385×. The tuned kernel saturated the L2 bandwidth at 77% of peak, indicating that further performance gains are unlikely without new hardware or fundamentally different approaches.

By Andrew James Amos
arXiv AI
Aug 18

Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps

arXiv:2608. 16309v1 Announce Type: cross Abstract: Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations and dynamic pruning mechanisms.

By Zirui Song, Yuye Zhu, Yang Yang