arXiv:2608. 05416v1 Announce Type: new Abstract: Can nonlinear dynamical systems be learned through a compact linear state-space representation, without directly solving a non-convex system-identification problem?
By Liane Galanti, Devan Shah, Shlomo Fortgang, Elad Hazan
The paper introduces recirculation, an inference‑time architectural enhancement for foundation models that reduces perplexity and improves accuracy on generation and reasoning tasks without adding significant latency. Recirculation adds a specific form of recurrence, enabling the model to function as a dynamical system that tracks belief states, and is distinct from chain‑of‑thought or depth‑recurrence methods. An adaptive variant requires minimal hyperparameter tuning and achieves notable gains on the Gemma3 family, including a 23% perplexity drop and a 21% accuracy increase on GSM8k.
By Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal, Rosanne Liu
arXiv:2605. 11287v2 Announce Type: replace-cross Abstract: A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transformers.
By Jevon Twitty, Vinh Pham, Nitiwith Rotchanarak, Viresh Pati, Yubin Kim, Shihao Yang, Jiecheng Lu
The paper introduces a framework for aligning the inductive bias of linear time‑invariant State Space Models (SSMs) with task‑specific spectral characteristics. By formalizing the bias through an SSM‑induced kernel and showing its spectrum is governed by the model’s frequency response, the authors propose Task‑Dependent Initialization (TDI), a fast power‑spectrum matching method. Experiments on synthetic data, one‑layer SSMs, and deep SSMs across real‑world benchmarks demonstrate that TDI improves data‑efficient generalization when the task’s spectral structure differs from the default SSM bias.
By Qiyu Chen, Guozhang Chen
arXiv:2602. 09530v2 Announce Type: replace-cross Abstract: We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization.
By Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael W. Mahoney
arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.
By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
arXiv:2608. 13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass.
By Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
By Shota Takashiro, Masanori Koyama, Takeru Miyato, Yusuke Iwasawa, Yutaka Matsuo, Kohei Hayashi
arXiv:2407. 07239v3 Announce Type: replace Abstract: Linear recurrent neural networks, such as State Space Models (SSMs) and Linear Recurrent Units (LRUs), have recently shown state-of-the-art performance on long sequence modelling benchmarks.
By Kai Biegun, Rares Dolga, Jake Cunningham, David Barber
Elastic Spectral State Space Models (ES-SSM) are a train‑once, export‑many sequence modeling framework that achieves elasticity by spectrally approximating the state‑space operator. The method builds on Hankel spectral filtering, using fixed spectral channels to represent long‑range token mixing and combining input‑adaptive gates with budget dropout to enable reliable deployment across different resource budgets. ES‑SSM is evaluated on byte‑level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning, showing that a single trained model can be truncated to competitive compact models while maintaining smooth quality‑cost curves across a wide range of truncation levels.
By Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
arXiv:2604. 18995v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction.
By Zhenbang Du, Kejing Xia, Xinrui Zhong, Yonggan Fu, Nicolai Oswald, Binfei Ji, Brucek Khailany, Pavlo Molchanov, Yingyan Lin
arXiv:2308. 13380v3 Announce Type: replace-cross Abstract: Is it possible to understand the intricacies of a dynamical system not solely from its input/output pattern, but also by observing the behavior of other systems within the same class?
By Marco Forgione, Filippo Pura, Dario Piga