arXiv AI By Mujtaba Farhan, Maheep Chaudhary

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning

Read the original on arXiv AI →

arXiv:2606. 07720v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
5d ago

Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization

The paper investigates whether different forms of intermediate computation in large language models—such as token-based traces, pause tokens, and latent reasoning—rely on the same underlying mechanism. By training five variants of GPTNeoX on an extended multi-hop reasoning task, the authors find that while vanilla, Chain-of-Thought, and Pause Token models perform well on in-distribution data, they fail to generalize to longer-hop out-of-distribution problems. In contrast, latent-reasoning models exhibit better depth generalization, with causal analysis revealing a sparse recurrent search circuit that implements forward reachability propagation across the graph.