arXiv Machine Learning

The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

arXiv:2605. 18079v2 Announce Type: replace Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconnect them from the models used in practice.

arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li
arXiv AI
Jun 19

Efficiently Representing Algorithms With Chain-of-Thought Transformers

arXiv:2606. 19697v1 Announce Type: cross Abstract: The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation.

By Yanhong Li, Anej Svete, Ashish Sabharwal, William Merrill
Hugging Face Trending Papers
Jun 17

Explaining Attention with Program Synthesis

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs.

arXiv Machine Learning
Jun 17

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

arXiv:2606. 17522v1 Announce Type: cross Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers.

By Vinoth Nandakumar, Qiang Qu, Pramod Thebe, Sakshi Khachariya, Tongliang Liu
arXiv Machine Learning
1d ago

Universal interpolation for deep residual self-attention networks

The paper proves that deep residual self‑attention networks can universally interpolate between any two collections of sequences using only two fixed single‑head attention blocks with Gaussian‑initialized projections. The interpolation is achieved by varying the order, signs, and durations of these blocks, independent of the specific input and output sequences. The result holds for both continuous and finite depth, and the authors also extend the analysis to causal‑masked settings.

By Sibylle Marcotte, Joan Bruna