The Token Before the Value Is the Key: How Hybrid Architectures Organize Induction Circuits
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates how to combine shared global key‑value (KV) caches with layer‑specific local history in decoder‑only Transformer language models. By separating historical content from the input source used to form it, the authors show that adding local history can reduce held‑out test perplexity by about 1.4% compared to a current‑token local branch, while also demonstrating benefits in capacity, entry‑count, and training‑compute controls. Experiments on a 126M‑parameter model with 2K context reveal that local history remains valuable even when adjacent layers share local inputs, and that a sufficient suffix schedule can reduce upper‑layer construction work without losing cache completeness.
Decoder-only Transformer language models cache keys and values (KV) to reuse past computation during generation. Sharing KV across layers saves storage but reduces the diversity of representations ava...
arXiv:2603.18908v5 Announce Type: replace Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
arXiv:2607. 15893v1 Announce Type: cross Abstract: While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising.
arXiv:2606. 15695v1 Announce Type: cross Abstract: Federated class-incremental learning (FCIL) becomes substantially harder when clients observe different label subsets, progress through tasks at different stages, and provide uneven supervision for the same semantic concepts.
arXiv:2609.36173v1 Announce Type: cross Abstract: Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuati...