arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.
By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
InfoMamba is an attention‑free hybrid model that combines a minimal‑bandwidth global interface with a selective recurrent stream. The architecture replaces token‑level self‑attention with a concept bottleneck linear filtering layer and integrates it via an information‑maximizing fusion (IMF) that injects global context into the state‑space dynamics. Experiments across classification, dense prediction, and non‑vision tasks show that InfoMamba outperforms strong Transformer and SSM baselines while maintaining near‑linear scaling and competitive accuracy‑efficiency trade‑offs.
By Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou
arXiv:2605. 11287v2 Announce Type: replace-cross Abstract: A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transformers.
By Jevon Twitty, Vinh Pham, Nitiwith Rotchanarak, Viresh Pati, Yubin Kim, Shihao Yang, Jiecheng Lu
arXiv:2606. 01306v1 Announce Type: new Abstract: While Transformer-based architectures have established themselves as a dominant paradigm in Multivariate Time Series Forecasting (MTSF), their core self-attention mechanism inherently functions as a low-pass filter, systematically smoothing out high-frequency signals vital for sharp local changes.
By Peng He, Yao Liu, Yanglei Gan, Run Lin, Yuxiang Cai, Qiao Liu
arXiv:2606. 23957v1 Announce Type: new Abstract: Learning Koopman operators with autoencoders enables linear prediction in a latent space, but long-horizon rollouts often drift off the learned manifold, leading to phase and amplitude errors on systems with switching, continuous spectra, or strong transients.
By Mohammed Nagdi, Evangelos-Marios Nikolados, Alexey Yermakov, Mars Gao, Nathan Kutz, Filippo Menolascina
arXiv:2609.39082v1 Announce Type: new
Abstract: As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and...
By Wentao Wang, Hengyu Zhong, Yunhan Jiang, Jialiang An, Meng Lu
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
arXiv:2605.26797v2 Announce Type: replace
Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden...
By Zeyi Huang, Xuehai He, LiLiang Ren, Yiping Wang, Baolin Peng, Hao Cheng, Shuohang Wang, Pengcheng He, Jianfeng Gao, Yong Jae Lee, Yelong Shen
The paper introduces the Progressive Memory Transformer (PMT), a transformer variant that adds writable, window‑aligned memory to expose mid‑range representations alongside token and sequence‑level outputs. PMT is trained with a hierarchical learning framework that applies separate objectives at local, mid‑range, and global scales, encouraging the model to capture fine‑grained variation, window‑level motifs, and overall sequence agreement. Experiments on seven UCR/UEA/UCI classification datasets, a cue‑retention probe, and forecasting tasks show that PMT achieves strong low‑label classification performance, competitive multi‑horizon forecasting, and evidence that its memory states encode mid‑range motifs.
By Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera
arXiv:2606. 27748v1 Announce Type: cross Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences.
By Haoran Zhang, Feng Zhou
arXiv:2607. 07706v1 Announce Type: new Abstract: The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference.
By Anna Kuzina, Paul N. Whatmough, Babak Ehteshami Bejnordi
The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.
By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi