Component-Weighted Centroid Search for Exact Incremental BPE
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.20276v2 Announce Type: replace-cross Abstract: Memoizing an expensive function of a sorted score vector is a data-structure problem before it is a numerical one: at a billion gridpoints, a...
arXiv:2609.13692v1 Announce Type: cross Abstract: LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definition...
The paper presents a table‑free index for tapered memoization grids, enabling compact out‑of‑core evaluation of functions that depend on sorted arguments. By showing that the grid’s key set corresponds to multiset combinations, the authors derive a closed‑form O(d) ranking and unranking scheme that removes the need for large preprocessing tables and allows order‑free parallel construction. The resulting values‑only flat array uses significantly less memory than hash‑map memoization, offers faster query times once cache limits are exceeded, and remains operable with memory‑mapped storage beyond RAM.
arXiv:2609.36766v1 Announce Type: new Abstract: Can a fixed continuous prefix replace a given low-rank adapter while the attention head stays frozen? In this research, we show that the answer depends...
arXiv:2606. 01502v1 Announce Type: cross Abstract: Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit is now a small, reusable chunk.
arXiv:2607. 14431v1 Announce Type: cross Abstract: We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights.