arXiv Machine Learning

Generalization Analysis of Transformers in Distribution Regression

arXiv:2606. 29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning.

arXiv Machine Learning
Aug 27

Cubit: Token Mixer with Kernel Ridge Regression

The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.

By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu
arXiv Machine Learning
Sep 10

Conditioned Initialization for Attention

arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...

By Hemanth Saratchandran, Simon Lucey
arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
arXiv Machine Learning
Jun 25

Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

arXiv:2606. 25010v1 Announce Type: new Abstract: Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learning are known to emerge abruptly past a certain model scale.

By Vatsal Baherwani, Zixi Chen, Shikai Qiu, Andrew Gordon Wilson, Pavel Izmailov
arXiv AI
Sep 25

Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs

The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.

By Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina
arXiv Machine Learning
Jun 3

Dynamic Short Convolutions Improve Transformers

arXiv:2606. 03825v1 Announce Type: new Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization.

By Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim