arXiv Machine Learning
Sep 21

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

The paper shows that multi‑hop retrieval failures cluster in predictable subpopulations and formalizes this with two theoretical results: (1) confident‑failure reduction is possible only when retrieval features carry mutual information about success, and (2) no single ANN score feature dominates across all failure regimes. Building on these insights, the authors introduce RegimeAbstain, which computes a Retrieval Confidence Score (RCS) from up to nine query‑ANN structural features and uses it to calibrate an abstention policy. Across three benchmarks and two retrieval architectures, RCS achieves the best or co‑best AUC‑AC and significantly reduces the Confident‑Wrong‑Answer Rate, demonstrating its effectiveness and domain‑agnostic applicability.

By Andre Bacellar
arXiv Machine Learning
3d ago

Relative Kinetic Utility: Calibrating Cross-Layer Credit for Global Structured LLM Pruning

The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.

By Tianhao Qian, Guilin Qi, Jiayu Chen