MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Lngram v2 introduces a latent N‑gram memory system that decouples memory routes, memory dimension, and backbone width, enabling scalable memory capacity for transformers. It employs context‑aware grouped‑query attention, a zero‑value sink, and counterfactual surrogate gradients to improve readout selectivity and routing trainability while preserving hard discrete addressing. Experiments on vision‑language models up to 30B parameters show consistent performance gains, reduced memory parameters, and stable semantic structure in the discrete IDs.
arXiv:2609.25537v1 Announce Type: new Abstract: Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing laten...
arXiv:2601. 07372v2 Announce Type: replace-cross Abstract: While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forcing them to inefficiently simulate retrieval through computation.
arXiv:2601. 21461v3 Announce Type: replace-cross Abstract: Modern sparse language models typically achieve sparsity through Mixture-of-Experts (MoE) layers, which dynamically route tokens to dense MLP "experts.
arXiv:2606. 10435v1 Announce Type: new Abstract: Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales quadratically with context length.
arXiv:2607. 21291v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost.