arXiv Computation and Language By Simon P. Villani

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

Read the original on arXiv Computation and Language →

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 7

What Attention Recalls and Recurrence Controls in Hybrid Language Models

Hybrid language models combine attention with a fixed-size recurrent state, yet the distinct roles of each component are not well understood. The authors introduce two cache-level interventions—split-prefill and state-swap—to isolate the contributions of the KV cache (attention) and the recurrent state. Experiments on Qwen3.5 and Falcon-H1 show that exact retrieval depends almost entirely on attention, while output language and persona rely mainly on recurrence, with the state-swap intervention confirming that answers derive from the KV side and language from the recurrent side.

By Kirill Afendulev, Alexey Dontsov, Elena Tutubalina, Anton Korznikov