arXiv:2605.23200v2 Announce Type: replace-cross
Abstract: The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate t...
By Junzhe Yang, Xiaoyu Shen
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
Neurosymbolics for Data Engineering introduces a neurosymbolic layer that can be added to existing LLM backbones to improve logical reasoning and reduce long‑context token usage. The layer boosts accuracy by an average of 85% on benchmarks such as BIRD‑CRITIC and LiveSQLBench without any task‑specific finetuning or RLHF. It also cuts effective token usage by over 50% and lowers time complexity from O(n²) to roughly O(n) for long‑context tasks.
By Vishvesh Bhat
The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.
By Maximilian Schambach, Clemens Biehl, Sam Thelin
arXiv:2609.25537v1 Announce Type: new
Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
By Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia, Yutaka Watanobe, Fang Liu
This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.
By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
arXiv:2601. 07372v2 Announce Type: replace-cross Abstract: While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forcing them to inefficiently simulate retrieval through computation.
By Xin Cheng, Rui Tian, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Chengqi Deng, Shangyan Zhou, Chenggang Zhao, Zhewen Hao, Yukun Li, Han Zhang, Zhengyan Zhang, Yixu Wei, M. Y Xu, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang
arXiv:2606. 27229v1 Announce Type: cross Abstract: Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the gate sees only the arriving token, not the memory it is about to modify.
By Sayak Dutta
The paper introduces Mahalanobis-Based Multi-Head Attention for Complex State Propagation (MHA‑CSP), a new attention mechanism that replaces the standard dot‑product with a Mahalanobis distance‑based RBF kernel. This approach enables infinite‑dimensional feature space attention without extra parameters, allows direct construction of Tree Attention via LogSumExp correction, and incorporates an attention meshing mechanism for cross‑head collaboration. Experiments show that with only 119K parameters and teacher forcing applied only at the final hidden state, MHA‑CSP outperforms Transformer and GCN baselines on long‑sequence state tracking tasks, demonstrating efficient structured reasoning.
By Xiaohe Li
Lngram v2 introduces a latent N‑gram memory system that decouples memory routes, memory dimension, and backbone width, enabling scalable memory capacity for transformers. It employs context‑aware grouped‑query attention, a zero‑value sink, and counterfactual surrogate gradients to improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments on vision‑language models up to 30B parameters show consistent performance gains, reduced memory parameters, and stable semantic structure in the discrete IDs.
By Yunao Zheng, Bin Wen, Xiaojie Wang
arXiv:2605. 18848v3 Announce Type: replace Abstract: This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by exploiting the exact decomposition property of kernel functions, thereby eliminating approximation error.
By Weinuo Ou
arXiv:2607. 21291v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost.
By Yidu Wu, Xiang Wang, Kejie Zhao, Zhangchi Wang, Qinghai Guo, Xiaoying Tang