arXiv:2606. 18694v1 Announce Type: new Abstract: A network of oscillators that synchronizes perfectly computes nothing further, so an attention architecture built from synchronization must locate its computation in structured departures from agreement.
By Joshua Nunley
arXiv:2606. 12059v1 Announce Type: new Abstract: We address transformer attention on energy-constrained physical substrates.
By Fabio Pasqualetti, Taosha Guo
The paper introduces a bounded spectral framework to analyze rotary attention in transformer language models, focusing on phase alignment, hidden‑state continuity, and semantic drift. It identifies ordered hidden‑state sequences as suitable domains for spectral decomposition, derives the Rotary Position Embedding (RoPE) attention score as a sum of magnitude‑weighted cosine terms, and proves a local stability lemma linking bounded phase displacement to pre‑softmax score degradation. By defining complex modal coordinates and a weighted coherence functional, the work distinguishes representational continuity from execution‑boundary admissibility, offering a theoretical program for when spectral structure explains continuity and when external governance is required.
By Abraham Chachamovits
arXiv:2609. 18145v1 Announce Type: new Abstract: Attention pays, at every layer and for every input, the cost of searching for whom to connect.
By Yoshiaki Takashita
arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).
By Luis Rosario Freytes
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
By Mark Oskin
The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.
By Barak Gahtan, Alex M. Bronstein
arXiv:2607. 09842v1 Announce Type: new Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes: no training (Gemma-4-E4B base), multimodal RLHF (Gemma-4-E4B-it), RL distillation (DeepSeek-R1-Distill-Qwen-7B), and SFT (Qwen2.
By Jorge A. Castillo, Marco Torres Y\'evenes, Juan Carlos Lanas
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
By Joshua S. Schiffman
arXiv:2607. 07478v1 Announce Type: new Abstract: FFT-based spectral preprocessing of learned query-key (Q/K) projections substantially improves transformer attention on character-level language modelling.
By Athanasios Zeris
arXiv:2605. 04893v2 Announce Type: replace Abstract: When a language model processes a hallucinated response, its attention routing tends to fail in one of two shapes: over-concentrating on a narrow set of positions, or spreading so diffusely that relevance is diluted, and the shape of the failure carries diagnostic signal.
By Dominik Dahlem, Diego Maniloff, Mac Misiura
arXiv:2609.28448v1 Announce Type: cross
Abstract: We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similar...
By Qucheng Gao, Zuyi Yang, Xiao Chen