arXiv Computation and Language

Prefix Sharing Is a Sorting Problem

arXiv AI
Aug 18

Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps

arXiv:2608. 16309v1 Announce Type: cross Abstract: Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single custom pipeline, leaving it unclear which findings transfer to modern engines with different index organizations and dynamic pruning mechanisms.

By Zirui Song, Yuye Zhu, Yang Yang
arXiv Machine Learning
1d ago

Local Search with Correlated Randomness

arXiv:2607.17469v2 Announce Type: replace-cross Abstract: How much does an algorithm's running-time distribution under independent randomness reveal about its behavior when independence is no longer...

By Yunbei Xu
arXiv Machine Learning
Sep 18

A Table-Free Index for Tapered Memoization Grids: Compact Out-of-Core Evaluation of Functions of Sorted Arguments

The paper presents a table‑free index for tapered memoization grids, enabling compact out‑of‑core evaluation of functions that depend on sorted arguments. By showing that the grid’s key set corresponds to multiset combinations, the authors derive a closed‑form O(d) ranking and unranking scheme that removes the need for large preprocessing tables and allows order‑free parallel construction. The resulting values‑only flat array uses significantly less memory than hash‑map memoization, offers faster query times once cache limits are exceeded, and remains operable with memory‑mapped storage beyond RAM.

By Tamal Maharaj
arXiv AI
Sep 23

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

The paper evaluates seven graph database engines, including Corvic AI, on a synthetic biomedical property graph with 1.02 million nodes and 5.34 million rows. It benchmarks query latency, bulk‑ingest throughput, point‑update latency, and correctness across a twenty‑query workload that covers neighborhood lookups, bounded paths, set intersections, anti‑joins, aggregation, ranking, temporal filters, full scans, and relational joins. The study finds that no single engine is universally fastest; performance depends on query shape, and the largest cost difference arises from bulk‑ingest throughput, which varies by three orders of magnitude and dominates total cost for workloads with fewer than about 10⁵ queries per data refresh.

By Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach
arXiv Machine Learning
Sep 7

Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving

The paper investigates how prefix caching, a default optimization in open‑source LLM serving stacks, affects reproducibility when combined with weight quantization. Experiments on an eighty‑episode multi‑turn agentic tool‑use workload show that enabling the cache causes the agent’s trajectory to change in 36.2 % of episodes at 16‑bit precision and 75.0 % at 4‑bit precision, while disabling the cache yields perfectly reproducible runs. The study identifies specific cache‑related settings that drive run‑to‑run divergence and demonstrates that cached serving is deterministic only when the cache state is preserved, which is not the case in typical deployments.

By Aditi Patodiya