The paper introduces OAttention, a token‑level attention mechanism that assigns each token a presence coefficient based on its hidden representation. This coefficient both gates the token’s output and weights its contribution to other tokens, making zero‑vector tokens behave as true zeros and enabling exact null‑receiver, null‑source, and empty‑support properties. The authors extend this idea to local O‑components and an O‑Transformer, and demonstrate small performance changes when retrofitting a pretrained TabPFN model.
By Heyang Gong
The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.
By Barak Gahtan, Alex M. Bronstein
arXiv:2605. 18848v3 Announce Type: replace Abstract: This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by exploiting the exact decomposition property of kernel functions, thereby eliminating approximation error.
By Weinuo Ou
arXiv:2606. 08105v1 Announce Type: new Abstract: When attention concentrates on a single token, a sink, what is the model actually computing?
By Lukas Fesser, Mozes Jacobs, Thomas Fel, Andy Keller, Sham Kakade
arXiv:2607. 23050v1 Announce Type: new Abstract: Neural scaling laws describe how loss decreases as models, data, and compute grow, but they do not answer a prior question: for a fixed task, what is the minimum model capacity required to solve it?
By Byeong Hoon Yoon
arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang
arXiv:2606. 20547v1 Announce Type: new Abstract: We place the attention token on the group: a token is an element $g_i$ of a matrix Lie group $G$ -- a bare transformation, with no feature payload and no external action $\rho(g)$ carrying it.
By Przemyslaw Musialski
The paper introduces Mahalanobis-Based Multi-Head Attention for Complex State Propagation (MHA‑CSP), a new attention mechanism that replaces the standard dot‑product with a Mahalanobis distance‑based RBF kernel. This approach enables infinite‑dimensional feature space attention without extra parameters, allows direct construction of Tree Attention via LogSumExp correction, and incorporates an attention meshing mechanism for cross‑head collaboration. Experiments show that with only 119K parameters and teacher forcing applied only at the final hidden state, MHA‑CSP outperforms Transformer and GCN baselines on long‑sequence state tracking tasks, demonstrating efficient structured reasoning.
By Xiaohe Li
arXiv:2609.01129v1 Announce Type: new
Abstract: We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under compos...
By Jiming Feng, Junliang Li
The paper introduces RefineICL, an attention‑gated, feed‑forward‑network‑free framework that refines representations in situ for tabular foundation models. By using support labels to guide episode‑specific updates, the method transfers learned corrections to unlabeled queries without altering model parameters, achieving state‑of‑the‑art performance on AMLB29 and TabArena benchmarks. Experiments and internal interventions demonstrate that intermediate support updates are essential for constructing task‑specific predictors in context.
By Tian Zhou, Beverly Jin, Linxiao Yang, Xue Wang, Wenwei Wang, Bingqing Peng, Mengni Ye, Jinjie Gu, Liang Sun
The projection of queries and keys are central to the attention mechanism in Transformer architectures. While they are mathematically symmetric, they play different roles in attention mechanisms.
The paper proposes a principled way to design hybrid transformer architectures that combine Full Attention (FA) and Linear Attention (LA). By introducing two intervention metrics—RoPE Frequency Importance Score (RFIS) and RoPE Positional Dependence (RPD)—the authors identify a clear taxonomy of retrieval and positional heads, defining a Global Positional Band (GPBand) that aligns with training-length positional scales. Using these insights, they build a Head‑wise Hybrid Architecture (HwH) that assigns FA to global retrieval and LA to local positional modeling, achieving strong language modeling, improved retrieval, and superior zero‑shot long‑context extrapolation compared to standard Transformers and other hybrids.
By Runlin Shi, Bojian Yin, Guoqi Li