arXiv Machine Learning By Wenbo Zhang, Xiang Ren

WhiteMatter: All-to-All Cross-Layer Connections via KV Mixing

Read the original on arXiv Machine Learning →

WhiteMatter introduces a novel architecture for Transformers that connects every attention layer to representations from all layers of each past token, allowing connection weights to vary across consumer layers and adapt to the source token. The design uses a router to mix the $L$ layer states of each token into $k$ KV channels, which are cached for subsequent tokens; each consumer layer attends to one channel. Experiments show that WhiteMatter outperforms a vanilla Transformer with 50% more layers and maintains most of this advantage even when the KV-cache is compressed by 50%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Sep 24

Shared Global KV with Layer-Specific Local History

The paper investigates how to combine shared global key‑value (KV) caches with layer‑specific local history in decoder‑only Transformer language models. By separating historical content from the input source used to form it, the authors show that adding local history can reduce held‑out test perplexity by about 1.4% compared to a current‑token local branch, while also demonstrating benefits in capacity, entry‑count, and training‑compute controls. Experiments on a 126M‑parameter model with 2K context reveal that local history remains valuable even when adjacent layers share local inputs, and that a sufficient suffix schedule can reduce upper‑layer construction work without losing cache completeness.

By Xinglang Xian