arXiv:2609.38149v1 Announce Type: new
Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for informa...
By Dor Tirosh, Ido Amos, Mor Geva
The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.
By Eduardo Novaes Hering
arXiv:2605.26797v2 Announce Type: replace
Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden...
By Zeyi Huang, Xuehai He, LiLiang Ren, Yiping Wang, Baolin Peng, Hao Cheng, Shuohang Wang, Pengcheng He, Jianfeng Gao, Yong Jae Lee, Yelong Shen
arXiv:2608. 05806v1 Announce Type: cross Abstract: While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning.
By Chang Shi, Tim Pearce, Manan Tomar, Siddhartha Sen, John Langford
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
arXiv:2609.36590v1 Announce Type: cross
Abstract: Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeof...
By Hankun Lin, Patrick Pynadath, Ruqi Zhang
arXiv:2609.17184v1 Announce Type: new
Abstract: Looped Transformers achieve strong performance with compact parameter sizes by repeatedly applying a shared stack of Transformer blocks across recurren...
By SangLyul Cho, Langqing Cui, Sehoon Kim, Dongsu Han, Insu Han
arXiv:2609.27233v1 Announce Type: new
Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...
By Zixuan Lan, Jessica Yang, Yanhong Li, Karen Livescu, Jiawei Zhou
arXiv:2511. 08577v3 Announce Type: replace-cross Abstract: Improving the reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applications.
By Tianyu Fu, Yichen You, Zekai Chen, Guohao Dai, Huazhong Yang, Yu Wang
arXiv:2610.01942v1 Announce Type: new
Abstract: Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of...
By Efstathios Karypidis, Spyros Gidaris, Nikos Komodakis
arXiv:2607. 15178v1 Announce Type: cross Abstract: Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time.
By Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He, Sanjeev Arora
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
By Shota Takashiro, Masanori Koyama, Takeru Miyato, Yusuke Iwasawa, Yutaka Matsuo, Kohei Hayashi