On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper investigates Associative Recall (AR) in the Mamba architecture, showing that Mamba implicitly learns linear hash functions to perform recall. It identifies the low‑level circuit responsible for this behavior and develops a theoretical framework—Recall Scaling Laws—based on similarity‑preserving hashing principles. The framework predicts embedding and state dimensions for perfect recall, recall success probability, and analyzes multi‑layer and multi‑head SSM patterns, with empirical results confirming its accuracy.
arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).
arXiv:2609.25802v1 Announce Type: new Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...
arXiv:2607. 13591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks.
arXiv:2602. 03315v2 Announce Type: replace Abstract: Agent memory systems must accommodate continuously growing information while supporting efficient, context-aware retrieval for downstream tasks.
arXiv:2601.12075v2 Announce Type: replace Abstract: Language models used in retrieval-augmented settings must arbitrate between parametric knowledge stored in their weights and contextual information...