Transformers converge to invariant algorithmic cores
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
The paper investigates whether a global workspace—a set of verbalisable, causally potent representations—emerges in transformer models that use recurrence instead of a stack of distinct layers. Using a Jacobian lens extended with a virtual‑unrolling adapter, the authors analyze two recurrent transformer architectures, Ouro‑2.6B and Huginn‑0125, and compare them to a standard Qwen3.6‑27B baseline. They find that a workspace does form in the iterated parts of both models, but recurrence alters how it can be accessed: Ouro reconstructs workspace content in every loop and requires writes and ablations across all loops, whereas Huginn forwards content across all recurrences but limits reads, writes, and ablations to a sliding window of about two recurrences.
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
arXiv:2604. 17121v3 Announce Type: replace Abstract: Transformers encode structure in sequences via an expanding contextual history.
The paper examines a method called Program-of-Layers (PoLar) that allows transformer layers to be dynamically routed rather than processed in a fixed sequence, mirroring the brain’s thalamic routing. Reproductions across five models confirm that skipping, repeating, and combining layer blocks improve performance, with shorter programs for easier inputs and more repeats for harder ones. However, the study could not replicate the claimed advantage of a learned single‑shot router, noting that its top prediction defaults to the standard pass while the top‑k predictions still yield accuracy gains. The authors also analyze the robustness of correction programs, finding them brittle to single edits, and release their code publicly.
arXiv:2608. 10251v1 Announce Type: cross Abstract: A transformer's answer lives on one axis: the direction its unembedding reads.
arXiv:2609.39892v1 Announce Type: new Abstract: Looped Transformers repeatedly apply the same set of Transformer layers, giving them a recurrent architecture for latent computation. Their strong perf...
arXiv:2606. 27538v1 Announce Type: cross Abstract: We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters the block.
The paper introduces recirculation, an inference‑time architectural enhancement for foundation models that reduces perplexity and improves accuracy on generation and reasoning tasks without adding significant latency. Recirculation adds a specific form of recurrence, enabling the model to function as a dynamical system that tracks belief states, and is distinct from chain‑of‑thought or depth‑recurrence methods. An adaptive variant requires minimal hyperparameter tuning and achieves notable gains on the Gemma3 family, including a 23% perplexity drop and a 21% accuracy increase on GSM8k.
arXiv:2608. 04879v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized.
arXiv:2603.21676v2 Announce Type: replace-cross Abstract: Standard Transformers have a fixed computational depth, limiting their ability to generalize to tasks that require variable-depth reasoning....
RecurTrace introduces adaptive latent reasoning for language models by allowing each looped layer to attend to its own past states and by using a halting head to decide when to stop iterating. This approach overcomes two limitations of prior latent recurrence methods: limited access to earlier computations and a fixed loop count that mismatches input difficulty. In experiments on MathQA, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, outperforming fixed‑depth baselines and other adaptive methods, and it also improves generation accuracy across a range of model sizes.
arXiv:2606. 00926v1 Announce Type: new Abstract: Mechanistic studies of sequence models often treat layerwise state encodings as architectural traits: recurrent models concentrate readable state, attention-based models distribute it.
arXiv:2607. 13491v1 Announce Type: cross Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters.