arXiv AI

Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers

arXiv:2608. 05472v1 Announce Type: cross Abstract: Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample operator mapping aggregated values to outputs is the same for every input set.

arXiv Machine Learning
1d ago

Universal interpolation for deep residual self-attention networks

The paper proves that deep residual self‑attention networks can universally interpolate between any two collections of sequences using only two fixed single‑head attention blocks with Gaussian‑initialized projections. The interpolation is achieved by varying the order, signs, and durations of these blocks, independent of the specific input and output sequences. The result holds for both continuous and finite depth, and the authors also extend the analysis to causal‑masked settings.

By Sibylle Marcotte, Joan Bruna
arXiv Machine Learning
Aug 27

Cubit: Token Mixer with Kernel Ridge Regression

The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.

By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu
arXiv AI
Aug 26

Mahalanobis-Based Multi-Head Attention for Complex State Propagation

The paper introduces Mahalanobis-Based Multi-Head Attention for Complex State Propagation (MHA‑CSP), a new attention mechanism that replaces the standard dot‑product with a Mahalanobis distance‑based RBF kernel. This approach enables infinite‑dimensional feature space attention without extra parameters, allows direct construction of Tree Attention via LogSumExp correction, and incorporates an attention meshing mechanism for cross‑head collaboration. Experiments show that with only 119K parameters and teacher forcing applied only at the final hidden state, MHA‑CSP outperforms Transformer and GCN baselines on long‑sequence state tracking tasks, demonstrating efficient structured reasoning.

By Xiaohe Li
arXiv Machine Learning
Sep 25

Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning

Transformers can learn broad families of tasks during pretraining and adapt to unseen tasks from a short prompt, but a rigorous understanding of this capability is limited. This paper studies how shared cross‑task structure influences the sample complexity of in‑context learning (ICL) by characterizing task‑space complexity through covering numbers, yielding a set of anchor functions that localize unseen tasks and predict responses. The authors construct a Transformer with Softmax attention to approximate this procedure and derive an error bound that separates the effects of pretraining tasks and prompt length, showing that once enough tasks are available the dependence on prompt length becomes dimension‑free.

By Zhongjie Shi, Rongjie Lai, Alexander Cloninger, Wenjing Liao
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen