arXiv:2609.28448v1 Announce Type: cross
Abstract: We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similar...
By Qucheng Gao, Zuyi Yang, Xiao Chen
arXiv:2609.36230v1 Announce Type: new
Abstract: We study the dynamical behavior of tokens in transformers from a control-theoretic perspective. Our model includes the feed-forward layer present after...
By Thomas Jacob Maranzatto, Semih Akkoc, Sennur Ulukus
Selective state space models (SSMs) use a recurrence to mix token information, a process analogous to attention in transformers. By modeling token evolution as an ordinary differential equation and applying input‑to‑state stability, the study proves that SSMs exhibit local exponential stability of consensus equilibria and delineates their domain of attraction for time‑varying weight matrices. Experiments on a pretrained Mamba‑2 model reveal that the output gate controls the degree of consensus, preventing tokens from fully converging.
By Jo\~ao Pedro Silvestre, \'Alvaro Rodr\'iguez Abella, Paulo Tabuada
arXiv:2609.24202v1 Announce Type: new
Abstract: Sparse attention reduces the quadratic cost of global self-attention while retaining strong empirical performance, but how its restricted interactions...
By Jingkun Liu, Yue Song
arXiv:2501. 18322v2 Announce Type: replace Abstract: Transformers, which are state-of-the-art in most machine learning tasks, represent the data as sequences of vectors called tokens.
By Val\'erie Castin, Pierre Ablin, Jos\'e Antonio Carrillo, Gabriel Peyr\'e
arXiv:2608. 18592v1 Announce Type: new Abstract: Whether distinct neural architectures develop common collective dynamics remains an open question.
By Byung Gyu Chae