arXiv:2607. 07066v1 Announce Type: cross Abstract: Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition.
By Zitong Andrew Chen, Junaid Hasan, Akhil Srinivasan, Hemkesh Bandi, Jarod Alper
arXiv:2603. 05556v2 Announce Type: replace Abstract: Integer sequences in the OEIS span values from single-digit constants to astronomical factorials and exponentials, making prediction challenging for standard tokenised models that cannot handle out-of-vocabulary values or exploit periodic arithmetic structure.
By Kazuhisa Nakasho
arXiv:2601. 22002v5 Announce Type: replace Abstract: Transformers achieve superior performance on many tasks, but impose heavy compute and memory requirements during inference.
By Anderson de Andrade, Alon Harell, Ivan V. Baji\'c
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
By Joshua S. Schiffman
arXiv:2604. 13082v2 Announce Type: replace-cross Abstract: Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the source of that delay remains poorly understood.
By Laura Gomezjurado Gonzalez
arXiv:2609.15975v1 Announce Type: cross
Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
By Shwai He, Haichao Zhang, Shen Yan
arXiv:2609.06000v1 Announce Type: cross
Abstract: We propose ModularPhaseNet, a classical and integer-computable discretization of the continuous complex phase geometry introduced in QuantumPhaseNet....
By Kiyotaka Kasubuchi, Kazuo Fukiya
arXiv:2607. 17843v1 Announce Type: new Abstract: Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles.
By Tongtian Zhu
The paper introduces the Von‑Neumann State‑Space Transformer (VN‑SST), a memory‑augmented Transformer that replaces the standard feed‑forward block with a low‑rank instruction bank. By decoding token‑specific operators from a low‑dimensional state‑space memory, VN‑SST achieves higher data‑efficiency and parameter‑efficiency on motor‑cortex neural‑decoding tasks and on small language‑model benchmarks. The model demonstrates that a compact instruction set can act as a control channel, improving performance without increasing accuracy through larger parameter counts.
By Morteza Sarafyazd
arXiv:2606. 00045v1 Announce Type: new Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact mathematical symmetries, such as modular arithmetic and non-commutative algebra.
By Sungyong Chung, Alireza Talebpour
arXiv:2608.30720v1 Announce Type: new
Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
By Kieran Murphy
arXiv:2607. 17166v1 Announce Type: new Abstract: Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks.
By Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen