arXiv AI

The Impossibility Triangle of Long-Context Modeling

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

arXiv AI
Jul 14

Extending LLM Context via Associative Recurrent Memory

arXiv:2607. 11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling.

By Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov
arXiv AI
Sep 10

Memory in Deep Time-Series Models

arXiv:2609.06006v1 Announce Type: cross Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state...

By Minh Hoang Nguyen, Huu Hiep Nguyen, Manh Nguyen, Van Dai Do, Dung Nguyen, Hung Le
arXiv Machine Learning
Sep 10

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.

By Yuval Koren, Assaf Ben-Kish, Raja Giryes, Lior Wolf, Itamar Zimerman
arXiv AI
Jul 14

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

arXiv:2607. 09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence.

By Siddharth Pal, Viktoria Rojkova
arXiv AI
2d ago

Reshape and Recur: Improving SSMs with Input Reshaping and Depth Recurrence

The paper proposes two extensions to State Space Models (SSMs) to reduce memory usage and improve performance. First, it introduces depth recurrence, allowing a looped SSM with fewer parameters to match the performance of a larger, non-recurrent model. Second, it advocates using a fixed time granularity across tasks by reshaping input sequences, which enhances how information is presented to the model. Both techniques consistently benefit four representative SSM architectures (LRU, S5, LinOSS, LrcSSM).

By M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu