arXiv Machine Learning

From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline

The paper reports a comprehensive tuning of both a novel sparse self‑organizing map algorithm (SparseBin) and its baseline cuSPARSE implementation. By optimizing four key levers—tile size, tile‑membership clustering, neuron‑axis chunking, and vectorised loads—the authors achieved a 5.6‑10.1× speed‑up per epoch for map sizes ranging from 32×32 to 512×512, and increased the performance margin over the CUDA baseline from ~80× to ~385×. The tuned kernel saturated the L2 bandwidth at 77% of peak, indicating that further performance gains are unlikely without new hardware or fundamentally different approaches.

arXiv Machine Learning
Aug 26

A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU

The paper presents a memory‑efficient sparse‑binary self‑organising map (SOM) that scales a MEDLINE atlas to over a million neurons on a single consumer GPU. By re‑ordering the codebook into a feature‑major layout, the authors accelerate the best‑matching‑unit search by 4.5–8.5× without increasing quantisation error, enabling training of a 1,048,576‑neuron SOM in 72 s on a 24 GB GPU. The approach outperforms existing cuSPARSE and CPU‑based SOM implementations, achieving the largest SOM reported to date and demonstrating that resolution limits are computational rather than data‑driven.

By Andrew James Amos
arXiv AI
Aug 26

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

The paper compares two strategies for handling memory limits in large language model (LLM) serving: tensor parallelism, which distributes weights and KV cache across multiple GPUs, and KV compression, which reduces cache size via quantisation and eviction on a single GPU. Using a cost‑normalised simulator calibrated on A100, A40, and H100 hardware, the authors find that across two models (Llama‑2 7B and 70B) and various GPU configurations, compression consistently outperforms tensor parallelism in cost per million tokens, offering 1.20× to 2.00× savings. The study identifies a model‑size threshold (~36B parameters on an 80 GB card) where compression dominates, while tensor parallelism becomes necessary only for larger models where weights alone exceed a single GPU’s capacity.

By Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal, Jay Iyer, Sai Pavan Kumar, Sai Kapil Kumar, Ramesh Nampelly, Raj Dandekar
arXiv AI
3d ago

ENAS: An Efficient Hardware-Aware Neural Architecture Search Framework for TinyML on Resource-Constrained Microcontrollers

ENAS is a hardware‑aware neural architecture search framework tailored for TinyML on microcontrollers. It uses a static feasibility check, a cell‑based search space with various block types and skip connections, and a three‑stage hybrid search strategy (random → top‑K → mutation) with cross‑run caching. The framework runs efficiently without GPUs, achieving significant search‑time speedups and competitive accuracy on Visual Wake Words and Melanoma Cancer benchmarks across a range of microcontrollers.

By Mohd Moin Khan, Naman Srivastava, Pandarasamy Arjunan
arXiv Machine Learning
Jul 7

Tile-Level Activation Overlap for Efficient LLM Inference

arXiv:2607. 02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.

By Abhinav Jangda, Tyler Sorensen, Sebastian Burckhardt, Jianlan YE, Chaoyin Li, Atul Gupta
arXiv Computation and Language
Sep 14

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

AMDKernelVault is an open HIP and Triton kernel corpus and training framework designed for AMD CDNA GPUs. It includes 62,153 verified HIP kernels, 39,893 Triton kernels, and 2,377 ROCm library QA entries, and introduces agent-driven pipelines (HIPKernelGen and TritonKernelGen) that convert PyTorch references into GPU kernels, compile, validate, and profile them on AMD hardware. The corpus was used to fine‑tune Qwen3-8B, achieving the highest correctness on several benchmarks such as PyTorch-to-HIP, TritonBench‑G, and ROCmBench under fixed evaluation budgets.

By Ji Liu, Saptarshi Majumder, Yiqing Huang, Wenwen Ouyang, Umang Pandey, Zeping Li, Chushi Chen, Zihao An, Puyuan Yang, Zekai Li, Sina Rafati, Ziqiong Liu, Pratik Prabhanjan Brahma, Dong Li, Zicheng Liu, Sharon Zhou, Emad Barsoum