When Recursive Models Finish Computing
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
Latent Recurrent Thoughts (LRT) proposes a method for reasoning with frozen large language models by operating in the model’s continuous representation space. A small auxiliary network generates initial latent vectors, which a tiny recurrent reasoner refines over multiple steps, decoupling computational depth from model size. Experiments on symbolic and natural‑language reasoning tasks show that LRT outperforms prior frozen‑decoder continuous‑space methods and chain‑of‑thought prompting while using far less inference compute.
arXiv:2609.33149v2 Announce Type: replace Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how...
Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic ta...
arXiv:2609.39967v1 Announce Type: cross Abstract: Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few...
Fast Weight Attention for Continual Learning introduces recurrent fast‑weight memories and selective state‑space models that compress expanding context into a fixed‑size recurrent state, enabling an online learning rule for state transitions. The paper derives normalized first‑order updates for squared‑error regression and negative inner‑product objectives, presenting several variants (Falcon‑1, Falcon‑2, Falcon‑3 and their inner‑product counterparts) with recurrent, masked‑parallel, and chunk‑parallel implementations. These methods demonstrate competitive performance in language modeling and improved length extrapolation on variable‑digit addition tasks.
RecurTrace introduces adaptive latent reasoning for language models by allowing each looped layer to attend to its own past states and by using a halting head to decide when to stop iterating. This approach overcomes two limitations of prior latent recurrence methods: limited access to earlier computations and a fixed loop count that mismatches input difficulty. In experiments on MathQA, RecurTrace achieves 56.9% accuracy with an average of 2.0 loops, outperforming fixed‑depth baselines and other adaptive methods, and it also improves generation accuracy across a range of model sizes.