The paper shows that multi‑hop retrieval failures cluster in predictable subpopulations and formalizes this with two theoretical results: (1) confident‑failure reduction is possible only when retrieval features carry mutual information about success, and (2) no single ANN score feature dominates across all failure regimes. Building on these insights, the authors introduce RegimeAbstain, which computes a Retrieval Confidence Score (RCS) from up to nine query‑ANN structural features and uses it to calibrate an abstention policy. Across three benchmarks and two retrieval architectures, RCS achieves the best or co‑best AUC‑AC and significantly reduces the Confident‑Wrong‑Answer Rate, demonstrating its effectiveness and domain‑agnostic applicability.
By Andre Bacellar
arXiv:2609.01556v1 Announce Type: cross
Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not...
By Nabira Rashid, Manolis Kellis
arXiv:2608. 15428v1 Announce Type: cross Abstract: Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question.
By Volodymyr Ovcharov
arXiv:2608.24079v1 Announce Type: cross
Abstract: A shared search-and-recommendation index must score new items from features alone because search has no exploration slot. In a public log covering bo...
By Theodore Rogers, Joe Standerfer, Dmitrii Timoshenko, Haoxue Li, Zuhaib Akhtar, Soyoung Yang
The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.
By Tanay Chowdhury, Saeideh Shahrokh Esfahani
The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking outperforms single‑action RL methods, while KG‑path selectors achieve near‑perfect unique‑output rates, and the overall end‑to‑end build cost is around $15 on commodity hardware.