arXiv Machine Learning By Hayden Helm, Ben Johnson, Carey Priebe

Query-efficient model evaluation using cached responses

Read the original on arXiv Machine Learning →

arXiv:2605. 07096v2 Announce Type: replace Abstract: Evaluating a new model on an existing benchmark is often necessary to understand its behavior before deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 19

Closing the Calibration Gap in Semantic Caching

arXiv:2606. 19719v1 Announce Type: cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries.

By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
arXiv Machine Learning
Aug 31

Closing the Operational Gap in Semantic Caching

Semantic caching reduces LLM inference costs by returning cached responses for semantically similar queries, but current evaluation using PR‑AUC only ranks scores and ignores usability at a fixed threshold, leading to poor deployment choices. The authors propose a cache‑aware metric, Precision–Cache Hit Ratio (P‑CHR) AUC, and an Operational Retention Rate (ORR) to measure how offline ranking quality translates to deployment. They decompose the operational gap into a recoverable threshold‑utility component and an irreducible structural component, showing that the gap is driven by the training objective rather than data scale and can be mitigated by score re‑normalization or objective changes, framing model selection as a threshold‑utility problem.

By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal