Google AI Blog By Google AI

Mixed-input matrix multiplication performance optimizations

Read the original on Google AI Blog →

Posted by Manish Gupta, Staff Software Engineer, Google Research AI-driven technologies are weaving themselves into the fabric of our daily routines, with the potential to enhance our access to knowledge and boost our overall productivity. The backbone of these applications lies in large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Google AI Blog.

arXiv AI
3d ago

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

The paper proposes a new matrix multiplication approach called SFC-CA GEMM that uses space‑filling curves to partition work in a platform‑ and shape‑oblivious way, achieving communication‑avoiding properties. It demonstrates provable asymptotic communication optimality for both square and rectangular matrices and outperforms vendor libraries on multiple x86 and Arm platforms, with speedups up to 5.5× for specific shapes and 1.8× in weighted harmonic mean throughput. The method is applied to real‑world tasks, improving large‑language‑model inference by up to 1.85× and distributed‑memory GEMM by up to 2.3× over state‑of‑the‑art frameworks.

By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
arXiv AI
Jul 14

CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning

arXiv:2512. 02551v3 Announce Type: replace-cross Abstract: In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels.

By Songqiao Su, Xiaoya Li, Albert Wang, Guoyin Wang, Jiwei Li, Chris Shum
arXiv Machine Learning
1d ago

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

The paper introduces TACO, a new optimizer for fine‑tuning large language models that drastically reduces optimizer state memory while preserving first‑order gradients. TACO selects the sign of the largest magnitude entry in each column of weight matrices, achieving a 174× reduction in persistent optimizer memory compared to AdamW8bit and a 2.9× decrease in peak training memory on OPT‑13B. This allows full‑parameter fine‑tuning of 30–32B‑parameter models on a single 80 GB GPU across multiple model families and tasks, with comparable accuracy and runtime to existing methods.

By Jichao Jiang (University of Central Florida), Cristian McGee (University of Central Florida), El Houcine Bergou (Mohammed VI Polytechnic University), Hanqin Cai (University of Central Florida), Aritra Dutta (University of Central Florida)
arXiv Machine Learning
Jun 30

DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures

arXiv:2511. 15503v3 Announce Type: replace-cross Abstract: High-performance Host processors can integrate Processing-In-Memory (PIM) devices, which can accelerate memory-intensive kernels of Machine Learning (ML) models, including Large Language Models (LLMs), by leveraging the large memory bandwidth available at PIM cores.

By Peiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati, Onur Mutlu, Gennady Pekhimenko, Christina Giannoula