Written as a Record, Read as an Address: What a Forward Pass Leaves in an Operation's KV Cache
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 27539v1 Announce Type: new Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record.
arXiv:2606. 17107v1 Announce Type: cross Abstract: Prefix caching reuses prefill only across an exactly shared prefix, so one changed field invalidates the entire downstream cache.
The paper introduces Galahad, a memory layer that stores a transformer language model’s key‑value state for blocks of text, allowing subsequent requests to reuse previously computed attention rather than recomputing it. On seven real‑world datasets, 98.7% of prompt tokens were already read, and with Galahad the model could attend to an entire 97,000‑token corpus, achieving 98–100% recall on a 100‑fact test while reducing inference time and energy consumption dramatically. The approach was validated across 30 models and all runtimes, demonstrating that stateful inference can replace stateless serving without loss of accuracy.
arXiv:2608. 11218v1 Announce Type: new Abstract: Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text.
A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out...
arXiv:2608. 11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC).