arXiv AI By Muhammad Mansoor, Tahir Ahmad, Yeo-Chan Yoon

Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs

Read the original on arXiv AI →

arXiv:2607. 04281v1 Announce Type: cross Abstract: Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar queries, but most existing methods do not model the time-varying freshness of open-web evidence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

Closing the Operational Gap in Semantic Caching

Semantic caching reduces LLM inference costs by returning cached responses for semantically similar queries, but current evaluation using PR‑AUC only ranks scores and ignores usability at a fixed threshold, leading to poor deployment choices. The authors propose a cache‑aware metric, Precision–Cache Hit Ratio (P‑CHR) AUC, and an Operational Retention Rate (ORR) to measure how offline ranking quality translates to deployment. They decompose the operational gap into a recoverable threshold‑utility component and an irreducible structural component, showing that the gap is driven by the training objective rather than data scale and can be mitigated by score re‑normalization or objective changes, framing model selection as a threshold‑utility problem.

By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
arXiv Machine Learning
Jun 19

Closing the Calibration Gap in Semantic Caching

arXiv:2606. 19719v1 Announce Type: cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries.

By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
arXiv AI
Sep 11

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall (FR) introduces an ontology-driven policy layer that categorizes personal facts into over ten behavioral types and applies tailored lifecycle rules—such as differential decay, supersession, and event-time validity—to manage memory persistence in large language models. The FR-Bank implementation, independent of underlying infrastructure, achieves a 76.9% pass rate on the new LifecycleBench benchmark and improves LongMemEval-S performance, while significantly reducing confabulation rates compared to prior systems. Ablation studies show that the generic lifecycle metadata drives correctness, whereas the behavioral ontology enhances calibration and reduces downstream hallucinations.

By Ansuman Mullick, Eray T\"uz\"un