arXiv AI

SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning

arXiv:2606. 26290v1 Announce Type: cross Abstract: While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accumulation remains under-explored.

arXiv Machine Learning
Jun 25

ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory

arXiv:2606. 25156v1 Announce Type: new Abstract: Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-value sequence grows, softmax probability mass can dilute across a wider distribution, inducing activation shift and long-context performance collapse.

By Habibullah Akbar
arXiv Machine Learning
Sep 18

Elastic Spectral State Space Models for Train-Once Budgeted Inference

Elastic Spectral State Space Models (ES-SSM) are a train‑once, export‑many sequence modeling framework that achieves elasticity by spectrally approximating the state‑space operator. The method builds on Hankel spectral filtering, using fixed spectral channels to represent long‑range token mixing and combining input‑adaptive gates with budget dropout to enable reliable deployment across different resource budgets. ES‑SSM is evaluated on byte‑level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning, showing that a single trained model can be truncated to competitive compact models while maintaining smooth quality‑cost curves across a wide range of truncation levels.

By Dachuan Song, Junyu Yin, Zechen Hu, Xuan Wang
arXiv AI
2d ago

Reshape and Recur: Improving SSMs with Input Reshaping and Depth Recurrence

The paper proposes two extensions to State Space Models (SSMs) to reduce memory usage and improve performance. First, it introduces depth recurrence, allowing a looped SSM with fewer parameters to match the performance of a larger, non-recurrent model. Second, it advocates using a fixed time granularity across tasks by reshaping input sequences, which enhances how information is presented to the model. Both techniques consistently benefit four representative SSM architectures (LRU, S5, LinOSS, LrcSSM).

By M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu
arXiv AI
Jun 10

Blurry Window Attention

arXiv:2606. 09862v1 Announce Type: cross Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form of KV cache, which becomes a bottleneck in long context scenarios.

By Axel Laborieux, Christos Sourmpis, Juan Gabriel Kostelec, Qinghai Guo
arXiv AI
Jun 4

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.

By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
arXiv AI
Jul 13

Self-Guided Test-Time Training for Long-Context LLMs

arXiv:2607. 09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs.

By Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu