Hugging Face Trending Papers

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active overwriting of stale ones. In our diagnostic suite, the strongest efficient baselines tend to solve only one side well.

arXiv Machine Learning
Aug 31

Fast Weight Attention for Continual Learning

Fast Weight Attention for Continual Learning introduces recurrent fast‑weight memories and selective state‑space models that compress expanding context into a fixed‑size recurrent state, enabling an online learning rule for state transitions. The paper derives normalized first‑order updates for squared‑error regression and negative inner‑product objectives, presenting several variants (Falcon‑1, Falcon‑2, Falcon‑3 and their inner‑product counterparts) with recurrent, masked‑parallel, and chunk‑parallel implementations. These methods demonstrate competitive performance in language modeling and improved length extrapolation on variable‑digit addition tasks.

By Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
arXiv Computation and Language
Sep 23

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

By Simon P. Villani
Hugging Face Trending Papers
Jun 29

Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management

We introduce Neural Subspace Reallocation (NSR), which reframes continual learning as memory management over parameter subspaces. Instead of treating Low-Rank Adaptation (LoRA) modules as disposable per-task adapters, NSR manages them as compressible, retrievable memory units on a frozen backbone through a recurring cycle: (1) compress learned LoRAs via SVD, (2) reserve them in a TaskKnowledgeBank, (3) recall related past LoRAs by embedding similarity to warm-start new or returning tasks, and (4) reallocate the active subspace accordingly, with distillation protecting prior tasks.

arXiv AI
Jul 14

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

arXiv:2607. 09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence.

By Siddharth Pal, Viktoria Rojkova
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa