arXiv AI By Yan Zhou

The Impossibility Triangle of Long-Context Modeling

Read the original on arXiv AI →

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

Extending LLM Context via Associative Recurrent Memory

arXiv:2607. 11614v1 Announce Type: cross Abstract: Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain constrained by quadratic compute and linear memory scaling.

By Gleb Kuzmin, Ivan Rodkin, Aydar Bulatov, Yuri Kuratov, Lyudmila Rvanova, Mikhail Katkov, Ilia Sochenkov, Misha Tsodyks, Timothy Baldwin, Mikhail Burtsev, Artem Shelmanov
arXiv AI
Sep 10

Memory in Deep Time-Series Models

arXiv:2609.06006v1 Announce Type: cross Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state...

By Minh Hoang Nguyen, Huu Hiep Nguyen, Manh Nguyen, Van Dai Do, Dung Nguyen, Hung Le
arXiv Machine Learning
Sep 10

On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.

By Yuval Koren, Assaf Ben-Kish, Raja Giryes, Lior Wolf, Itamar Zimerman