arXiv Machine Learning By Qingbo Wu, Ke Li, Wenzhu Wang, Jie Yu, Ruian Zhang, Lili Liu

Operator Fusion for LLM Inference on the Tensix Architecture

Read the original on arXiv Machine Learning →

arXiv:2606. 09879v1 Announce Type: new Abstract: This study addresses on-device inference bottlenecks of Transformer models on Tenstorrent's Tensix architecture and proposes an operator fusion strategy that enhances data locality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 18

MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators

MeshKV is a Network‑on‑Chip key‑value cache fabric designed to improve transformer decoding on tiled accelerators. It uses packetized flows, TaKV affine striping, Mare multicast with duplicate suppression, and Pad to overlap prefetch, multiply, and softmax operations, thereby converting back‑pressure into useful KV transfer. In an 8×8 FPGA prototype with LLaMA‑2‑7B and Mistral‑7B, MeshKV cuts interconnect traffic by up to 58 %, doubles KV bandwidth utilization, and boosts multi‑stream throughput by up to 1.9×.

By Dong Liu, Yanxuan Yu
arXiv AI
Aug 28

Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models

The paper reports the development of 2‑bit microkernels for CPUs and mixed‑precision 2‑bit kernels for Intel Xe2 GPUs, achieving near‑roofline performance. Integrated into LLM inference pipelines, these kernels deliver up to 7× speedup over 16‑bit inference on CPUs and 6.7× on GPUs, surpassing the current state‑of‑the‑art bitnet.cpp runtime by 2.2×. The work demonstrates that ultra‑low‑bit LLM models can be deployed efficiently, offering significant gains in latency, memory, throughput, and energy consumption.

By Evangelos Georganas, Dhiraj Kalamkar, Alexander Heinecke, Pradeep Dubey
arXiv AI
Sep 15

mKernel: Fast Multi-GPU, Multi-Node Fused Kernels

arXiv:2609.13585v1 Announce Type: cross Abstract: Communication has become a bottleneck in distributed training and inference of large models. Overlapping communication with computation at the granul...

By Ziming Mao, Yihan Zhang, Shawn Wei Chew, Shuang Ma, Costin Raiciu, Yang Zhou, Scott Shenker, Ion Stoica