arXiv Machine Learning

Closing the Calibration Gap in Semantic Caching

arXiv:2606. 19719v1 Announce Type: cross Abstract: Semantic caching cuts LLM inference costs by serving a cached response to semantically similar queries.

arXiv Machine Learning
Aug 31

Closing the Operational Gap in Semantic Caching

Semantic caching reduces LLM inference costs by returning cached responses for semantically similar queries, but current evaluation using PR‑AUC only ranks scores and ignores usability at a fixed threshold, leading to poor deployment choices. The authors propose a cache‑aware metric, Precision–Cache Hit Ratio (P‑CHR) AUC, and an Operational Retention Rate (ORR) to measure how offline ranking quality translates to deployment. They decompose the operational gap into a recoverable threshold‑utility component and an irreducible structural component, showing that the gap is driven by the training objective rather than data scale and can be mitigated by score re‑normalization or objective changes, framing model selection as a threshold‑utility problem.

By Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev, Srijith Rajamohan, Jen Agarwal
arXiv AI
Sep 3

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

The paper introduces SCX Router, a lightweight GLiClass-based model selector that assigns suitability scores to inference-time language models without autoregressive generation. It uses a 0.6B-parameter Qwen3 decoder with a shallow bidirectional scorer, preserving a text-only key–value cache across sessions and predicting task attributes such as type, difficulty, and expected output length. The authors build a comprehensive task ontology with 23 families, 115 types, and 1,173 synthetic examples, generating 150,000 verifier-scored tasks to train the router, which outperforms baseline models on LiveBench subsets with a top‑1 score of 0.707 versus 0.696 for the strongest fixed model.

By Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov
arXiv AI
Aug 25

CacheSpec: Finding the Sweet Spot for Small Models in Large Language Models

CacheSpec is an inference optimization framework that transforms Program-of-Thoughts (PoT) style programs into reusable cache objects for large language models. By employing a small model for semantic variable extraction on cache hits and speculative drafting during target-LLM generation, CacheSpec reduces inference latency and improves cache reuse. Experiments on shopping, web, formula, and code QA datasets demonstrate up to 3.1× speedup in latency and 2.8× throughput gains over traditional PoT methods, while maintaining or improving task quality.

By Jingquan Chen, Jie Feng, Jinghua Piao, Shaogang Hu, Yong Li