arXiv AI By Hanzhi Zhang, Qiao Zhang, Qinglei Cao, Heng Fan, Yan Huang, Kewei Sha, Yunhe Feng

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

Read the original on arXiv AI →

TileMix is a tile‑centric mixed‑precision attention kernel that routes score‑tile groups within fused dense attention to either FP16 or INT8 computation, using compact bitmasks to decide precision per tile. By partitioning the attention matrix into hardware‑aligned tiles and updating a shared online‑softmax state, TileMix preserves dense token connectivity without requiring training and supports grouped‑query attention, variable‑length batches, and INT8 key/value caches. Benchmarks on LLaMA, Qwen, and Vicuna show that TileMix restores long‑context quality lost with uniform INT8 and improves prefill throughput over FP16, offering a controllable accuracy‑efficiency trade‑off across model families.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 23

SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture

The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.

By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni
arXiv AI
Jun 12

MiniMax Sparse Attention

arXiv:2606. 13392v1 Announce Type: new Abstract: Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale.

By Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao