arXiv Machine Learning

Elastic Spectral State Space Models for Train-Once Budgeted Inference

Elastic Spectral State Space Models (ES-SSM) are a train‑once, export‑many sequence modeling framework that achieves elasticity by spectrally approximating the state‑space operator. The method builds on Hankel spectral filtering, using fixed spectral channels to represent long‑range token mixing and combining input‑adaptive gates with budget dropout to enable reliable deployment across different resource budgets. ES‑SSM is evaluated on byte‑level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning, showing that a single trained model can be truncated to competitive compact models while maintaining smooth quality‑cost curves across a wide range of truncation levels.

arXiv AI
Jun 9

End-to-End Context Compression at Scale

arXiv:2606. 09659v1 Announce Type: cross Abstract: Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length.

By Ang Li, Sean McLeish, Haozhe Chen, Nimit Kalra, Zaiqian Chen, Artem Gazizov, Venkata Anoop Suhas Kumar Morisetty, Bhavya Kailkhura, Harshitha Menon, Zhuang Liu, Brian R. Bartoldson, Tom Goldstein, Sanae Lotfi, Micah Goldblum, Pavel Izmailov
arXiv Machine Learning
Jun 25

Frequency Domain Reservoir Computing

arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.

By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
arXiv Machine Learning
Aug 14

A Simple State Space Model Excels at Multivariate Time Series Classification

arXiv:2605. 27406v2 Announce Type: replace Abstract: Structured state space models (SSMs) have recently emerged as a promising foundation for sequence modeling, with Mamba-based architectures demonstrating strong performance through input-dependent state transitions, albeit at considerable complexity.

By Hassan Saadatmand, Geoffrey I. Webb, Hamid Rezatofighi, Mahsa Salehi
arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
1d ago

Aligning Inductive Bias for Data-Efficient Generalization in State Space Models

The paper introduces a framework for aligning the inductive bias of linear time‑invariant State Space Models (SSMs) with task‑specific spectral characteristics. By formalizing the bias through an SSM‑induced kernel and showing its spectrum is governed by the model’s frequency response, the authors propose Task‑Dependent Initialization (TDI), a fast power‑spectrum matching method. Experiments on synthetic data, one‑layer SSMs, and deep SSMs across real‑world benchmarks demonstrate that TDI improves data‑efficient generalization when the task’s spectral structure differs from the default SSM bias.

By Qiyu Chen, Guozhang Chen
Hugging Face Trending Papers
Jun 8

End-to-End Context Compression at Scale

Long-context language model inference is bottlenecked by memory, as the KV cache grows with context length. Recent techniques to compress the KV cache fall short: they either degrade model quality substantially or require considerable time and compute to compress a single long prompt.

arXiv Machine Learning
Jul 14

Controllably Efficient Language Models

arXiv:2511. 05313v2 Announce Type: replace Abstract: The substantial inference costs of attention in transformers motivated the development of efficient sequence mixers: namely sparse and sliding window attention, convolutions and linear attention.

By Jatin Prakash, Aahlad Puli, Rajesh Ranganath
arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi