Introduction to Transformers: an NLP Perspective
arXiv:2311. 17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing.
Related stories
Training and Finetuning Embedding Models with Sentence Transformers
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.
Generalization Analysis of Transformers in Distribution Regression
arXiv:2606. 29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning.
Training and Finetuning Sparse Embedding Models with Sentence Transformers
A Survey of Transformer-based Language Models with Focus on Efficiency
The paper surveys Transformer-based large language models (LLMs) with a focus on efficiency, reviewing 312 articles that cover data curation, model design, downsizing, and dynamic inference. It also examines efficiency in adaptation strategies such as pre‑training, fine‑tuning, prompt‑engineering, and Retrieval‑Augmented Generation (RAG). A statistical analysis and evaluation of over 30 prominent NLP models on 13 benchmarks provide insights into both efficiency and efficacy, highlighting trends toward sustainable NLP practices.
Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects
The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
arXiv:2607. 17624v1 Announce Type: new Abstract: Transformers are remarkably versatile and their design is largely consistent across a variety of applications.
Transfer learning for conflict and duplicate detection in software requirement pairs
arXiv:2301. 03709v3 Announce Type: replace-cross Abstract: Consistent and holistic expression of software requirements is important for the success of software projects.
An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars
arXiv:2606. 17522v1 Announce Type: cross Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers.