arXiv Machine Learning By Andre Bacellar

Per-Query Gating of LLM Rerankers for Multi-Hop Retrieval

Read the original on arXiv Machine Learning →

arXiv:2609. 22880v1 Announce Type: cross Abstract: LLM rerankers add of the order of \$0.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 21

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

The paper shows that multi‑hop retrieval failures cluster in predictable subpopulations and formalizes this with two theoretical results: (1) confident‑failure reduction is possible only when retrieval features carry mutual information about success, and (2) no single ANN score feature dominates across all failure regimes. Building on these insights, the authors introduce RegimeAbstain, which computes a Retrieval Confidence Score (RCS) from up to nine query‑ANN structural features and uses it to calibrate an abstention policy. Across three benchmarks and two retrieval architectures, RCS achieves the best or co‑best AUC‑AC and significantly reduces the Confident‑Wrong‑Answer Rate, demonstrating its effectiveness and domain‑agnostic applicability.

By Andre Bacellar
arXiv AI
Aug 20

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.

By Tanay Chowdhury, Saeideh Shahrokh Esfahani
Hugging Face Trending Papers
Aug 19

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking outperforms single‑action RL methods, while KG‑path selectors achieve near‑perfect unique‑output rates, and the overall end‑to‑end build cost is around $15 on commodity hardware.