arXiv Machine Learning

WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing

WhiteMatter introduces a novel architecture for Transformers that connects every attention layer to representations from all layers of each past token, allowing connection weights to vary across consumer layers and adapt to the source token. The design uses a router to mix the $L$ layer states of each token into $k$ KV channels, which are cached for subsequent tokens; each consumer layer attends to one channel. Experiments show that WhiteMatter outperforms a vanilla Transformer with 50% more layers and maintains most of this advantage even when the KV-cache is compressed by 50%.

arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Sep 24

Shared Global KV with Layer-Specific Local History

The paper investigates how to combine shared global key‑value (KV) caches with layer‑specific local history in decoder‑only Transformer language models. By separating historical content from the input source used to form it, the authors show that adding local history can reduce held‑out test perplexity by about 1.4% compared to a current‑token local branch, while also demonstrating benefits in capacity, entry‑count, and training‑compute controls. Experiments on a 126M‑parameter model with 2K context reveal that local history remains valuable even when adjacent layers share local inputs, and that a sufficient suffix schedule can reduce upper‑layer construction work without losing cache completeness.

By Xinglang Xian
arXiv AI
Jul 8

DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression

arXiv:2607. 06523v1 Announce Type: new Abstract: Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet existing compression methods often apply uniform budgets across layers or tokens and degrade retrieval when lexical cues and semantic states require different preservation.

By Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera