arXiv Machine Learning

SpectraLDS: Provable Distillation for Linear Dynamical Systems

arXiv:2505. 17868v2 Announce Type: replace Abstract: We present the first provable method for identifying symmetric linear dynamical systems (LDS) with accuracy guarantees that are independent of the systems' state dimension or effective memory.

arXiv Machine Learning
Aug 19

Recirculation

The paper introduces recirculation, an inference‑time architectural enhancement for foundation models that reduces perplexity and improves accuracy on generation and reasoning tasks without adding significant latency. Recirculation adds a specific form of recurrence, enabling the model to function as a dynamical system that tracks belief states, and is distinct from chain‑of‑thought or depth‑recurrence methods. An adaptive variant requires minimal hyperparameter tuning and achieves notable gains on the Gemma3 family, including a 23% perplexity drop and a 21% accuracy increase on GSM8k.

By Michael C. Mozer, Shoaib Ahmed Siddiqui, Danny Sawyer, Sunny Sanyal, Rosanne Liu
arXiv Machine Learning
1d ago

Aligning Inductive Bias for Data-Efficient Generalization in State Space Models

The paper introduces a framework for aligning the inductive bias of linear time‑invariant State Space Models (SSMs) with task‑specific spectral characteristics. By formalizing the bias through an SSM‑induced kernel and showing its spectrum is governed by the model’s frequency response, the authors propose Task‑Dependent Initialization (TDI), a fast power‑spectrum matching method. Experiments on synthetic data, one‑layer SSMs, and deep SSMs across real‑world benchmarks demonstrate that TDI improves data‑efficient generalization when the task’s spectral structure differs from the default SSM bias.

By Qiyu Chen, Guozhang Chen
arXiv AI
Jul 7

Learning to Discover Iterative Spectral Algorithms

arXiv:2602. 09530v2 Announce Type: replace-cross Abstract: We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization.

By Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael W. Mahoney
arXiv Machine Learning
Jun 25

Frequency Domain Reservoir Computing

arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.

By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
arXiv Machine Learning
Jun 25

RotRNN: Modelling Long Sequences with Rotations

arXiv:2407. 07239v3 Announce Type: replace Abstract: Linear recurrent neural networks, such as State Space Models (SSMs) and Linear Recurrent Units (LRUs), have recently shown state-of-the-art performance on long sequence modelling benchmarks.

By Kai Biegun, Rares Dolga, Jake Cunningham, David Barber
arXiv Machine Learning
Sep 18

Elastic Spectral State Space Models for Train-Once Budgeted Inference

Elastic Spectral State Space Models (ES-SSM) are a train‑once, export‑many sequence modeling framework that achieves elasticity by spectrally approximating the state‑space operator. The method builds on Hankel spectral filtering, using fixed spectral channels to represent long‑range token mixing and combining input‑adaptive gates with budget dropout to enable reliable deployment across different resource budgets. ES‑SSM is evaluated on byte‑level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning, showing that a single trained model can be truncated to competitive compact models while maintaining smooth quality‑cost curves across a wide range of truncation levels.

By Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
arXiv AI
Jun 3

$R^2$-dLLM: Accelerating Diffusion Large Language Models via Spatio-Temporal Redundancy Reduction

arXiv:2604. 18995v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive generation by enabling parallel token prediction.

By Zhenbang Du, Kejing Xia, Xinrui Zhong, Yonggan Fu, Nicolai Oswald, Binfei Ji, Brucek Khailany, Pavlo Molchanov, Yingyan Lin