Semantic caching reduces LLM inference costs by returning cached responses for semantically similar queries, but current evaluation using PR‑AUC only ranks scores and ignores usability at a fixed threshold, leading to poor deployment choices. The authors propose a cache‑aware metric, Precision–Cache Hit Ratio (P‑CHR) AUC, and an Operational Retention Rate (ORR) to measure how offline ranking quality translates to deployment. They decompose the operational gap into a recoverable threshold‑utility component and an irreducible structural component, showing that the gap is driven by the training objective rather than data scale and can be mitigated by score re‑normalization or objective changes, framing model selection as a threshold‑utility problem.
By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
arXiv:2606. 19719v1 Announce Type: cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries.
By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
Current LLM memory systems treat all personal facts identically, so stores grow without bound while retrieval precision degrades. The core challenge is lifecycle management: which memories should pers...
Fortunate Recall (FR) introduces an ontology-driven policy layer that categorizes personal facts into over ten behavioral types and applies tailored lifecycle rules—such as differential decay, supersession, and event-time validity—to manage memory persistence in large language models. The FR-Bank implementation, independent of underlying infrastructure, achieves a 76.9% pass rate on the new LifecycleBench benchmark and improves LongMemEval-S performance, while significantly reducing confabulation rates compared to prior systems. Ablation studies show that the generic lifecycle metadata drives correctness, whereas the behavioral ontology enhances calibration and reduces downstream hallucinations.
By Ansuman Mullick, Eray T\"uz\"un
arXiv:2609.35908v1 Announce Type: cross
Abstract: Semantic caches reduce LLM serving costs by reusing previously generated answers for semantically similar queries. However, retrieval is based solely...
By Zihan Zhang, Shuangjie Yao, Zesen Liu, Zhixiang Zhang, Wai Ip Lai, Dung Hiu Hilton Yeung, Chun Kit Zhang, Fuchen Ma, Yuanyuan Yuan, Yu Jiang, Dongdong She
arXiv:2606. 26511v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) gives agents access to accumulated knowledge, but has no model of time.
By Neeraj Yadav
arXiv:2608. 07954v1 Announce Type: new Abstract: Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such as Think-on-Graph and Reasoning-on-Graph repeatedly query the same graph neighborhoods across different questions.
By Uros Stanic, Changcheng Yuan, Sabuj Laskar, Ariful Azad
arXiv:2608.30647v1 Announce Type: cross
Abstract: Language models can answer from precomputed memory, a model's saved reading of a body of material, reused across requests instead of read again at ea...
By Asa Shepard
arXiv:2606. 02581v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) faces a fundamental three-way tension: deeper retrieval improves factual grounding but inflates token costs and end-to-end latency.
By Sanjay Mishra
arXiv:2607. 15516v1 Announce Type: cross Abstract: Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt compression (fewer tokens sent).
By Yan Song
Clarify-Then-Search is a benchmark that tests whether large language models can ask clarification questions to improve the usefulness of deep search results. It uses 518 real-world query pairs from Baidu, where each intent query is paired with an underspecified version. The evaluation involves a clarifier asking up to three questions, a user answerer providing only explicit information, and a rewriter generating a new query that is then searched; performance is measured by a weighted nugget-recall score.
By Deqiang Huang, Jingbo Zhou, Xinjiang Lu, Tong Xu, Hua Wu, Enhong Chen
CacheWeaver is a lightweight prompt‑layer technique that orders evidence for Retrieval‑Augmented Generation (RAG) to improve cache reuse in serving engines like vLLM. By maintaining a prefix tree of recently served evidence sequences and greedily placing the most reusable prefix first, it reduces median time‑to‑first‑token by 20‑33 % across three vLLM configurations without harming answer quality. The greedy policy achieves 97.5 % of the gain possible with oracle ordering, showing that most reusable prefix locality can be recovered with a simple scheduling layer.
By Kaizhen Tan, Rong Gu, Mingyuan Li