Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost
Read the original on arXiv AI →The paper introduces Galahad, a memory layer that stores a transformer language model’s key‑value state for blocks of text, allowing subsequent requests to reuse previously computed attention rather than recomputing it. On seven real‑world datasets, 98.7% of prompt tokens were already read, and with Galahad the model could attend to an entire 97,000‑token corpus, achieving 98–100% recall on a 100‑fact test while reducing inference time and energy consumption dramatically. The approach was validated across 30 models and all runtimes, demonstrating that stateful inference can replace stateless serving without loss of accuracy.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.