MuonSSM: Orthogonalizing State Space Models for Sequence Modeling
arXiv:2606. 30461v1 Announce Type: new Abstract: State space models (SSMs) have emerged as efficient linear-time alternatives to attention for long-sequence modeling.
arXiv:2606. 26290v1 Announce Type: cross Abstract: While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accumulation remains under-explored.
arXiv:2606. 30461v1 Announce Type: new Abstract: State space models (SSMs) have emerged as efficient linear-time alternatives to attention for long-sequence modeling.
arXiv:2606. 16093v1 Announce Type: cross Abstract: Modeling long-range dependencies remains a central challenge in natural language processing.
arXiv:2606. 25156v1 Announce Type: new Abstract: Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-value sequence grows, softmax probability mass can dilute across a wider distribution, inducing activation shift and long-context performance collapse.
arXiv:2606. 12364v1 Announce Type: new Abstract: Transformers dominate modern sequence modeling, but their quadratic attention incurs substantial computational cost.
Elastic Spectral State Space Models (ES-SSM) are a train‑once, export‑many sequence modeling framework that achieves elasticity by spectrally approximating the state‑space operator. The method builds on Hankel spectral filtering, using fixed spectral channels to represent long‑range token mixing and combining input‑adaptive gates with budget dropout to enable reliable deployment across different resource budgets. ES‑SSM is evaluated on byte‑level language modeling, Long Range Arena, Speech Commands V2, and offline reinforcement learning, showing that a single trained model can be truncated to competitive compact models while maintaining smooth quality‑cost curves across a wide range of truncation levels.
arXiv:2606. 25156v3 Announce Type: replace-cross Abstract: Native length extrapolation remain a weakly solvable problem in language modeling due to trade-off balancing between exact retrieval fidelity, long-document likelihood, and inference efficiency.
The paper proposes two extensions to State Space Models (SSMs) to reduce memory usage and improve performance. First, it introduces depth recurrence, allowing a looped SSM with fewer parameters to match the performance of a larger, non-recurrent model. Second, it advocates using a fixed time granularity across tasks by reshaping input sequences, which enhances how information is presented to the model. Both techniques consistently benefit four representative SSM architectures (LRU, S5, LinOSS, LrcSSM).
arXiv:2605.07111v3 Announce Type: replace-cross Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
arXiv:2606. 09862v1 Announce Type: cross Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form of KV cache, which becomes a bottleneck in long context scenarios.
arXiv:2606. 24650v1 Announce Type: cross Abstract: We present Harmonic, a hierarchical state space model (SSM) for language modeling.
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
arXiv:2607. 09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs.