The paper shows that multi‑hop retrieval failures cluster in predictable subpopulations and formalizes this with two theoretical results: (1) confident‑failure reduction is possible only when retrieval features carry mutual information about success, and (2) no single ANN score feature dominates across all failure regimes. Building on these insights, the authors introduce RegimeAbstain, which computes a Retrieval Confidence Score (RCS) from up to nine query‑ANN structural features and uses it to calibrate an abstention policy. Across three benchmarks and two retrieval architectures, RCS achieves the best or co‑best AUC‑AC and significantly reduces the Confident‑Wrong‑Answer Rate, demonstrating its effectiveness and domain‑agnostic applicability.
By Andre Bacellar
arXiv:2608. 00585v1 Announce Type: cross Abstract: Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail.
By Randhir Kumar
arXiv:2606. 28367v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross-document summarization, graph-based expansion, per-query routing, rank fusion, and corrective re-retrieval.
By Sadanand Singh, Allam Reddy, Manan Chopra
arXiv:2609. 22880v1 Announce Type: cross Abstract: LLM rerankers add of the order of \$0.
By Andre Bacellar
arXiv:2608. 05153v1 Announce Type: cross Abstract: GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound.
By Meftun Akarsu, Burak Ozdemir
The paper introduces Global Relative Kinetic Utility (Global RKU), a label‑free method for calibrating cross‑layer credit in global structured pruning of large language models. Global RKU estimates channel importance via a final‑hidden‑state activation‑gradient signal and applies block‑relative normalization to remove block‑common scale while preserving within‑block ordering, enabling a single‑stage static pruning topology. Experiments on Qwen‑2.5‑7B show significant performance gains at various sparsity levels, and ablation studies confirm the effectiveness of the relative‑normalization step.
By Tianhao Qian, Guilin Qi, Jiayu Chen
arXiv:2609.30738v1 Announce Type: new
Abstract: KV cache eviction methods such as SnapKV and PyramidKV rank tokens solely by mean attention over a small observation window. We study a unified score,...
By Tianfang Xie, Wei Zhu
We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation through a multi-modal and personalized conversational recomme...
FCPRAG introduces a fusion-controlled parametric retrieval‑augmented generation framework that uses a lightweight controller to predict per‑passage fusion scores and sample‑level calibration signals, such as a mixing gate and adaptive temperature. This approach mitigates the bottleneck of evidence‑level fusion when multiple passages are retrieved, enabling selective fusion under informative signals and conservative fusion under uncertainty. Experiments on HotpotQA, 2WikiMultiHopQA, PopQA, and ComplexWebQuestions demonstrate consistent F1 improvements over standard RAG and parametric RAG baselines, with gains up to 4.65% on 2WikiMultiHopQA and 7.55% on CWQ, while also reducing tuning cost and enhancing robustness to retrieval perturbations.
By Jinchang Zhu, Jindong Li, Yi Ding, Xiaojian Nie, Rong Fu, Shuangyong Song, Haowei He, Menglin Yang
arXiv:2609.22770v2 Announce Type: cross
Abstract: We study a parameterized hybrid ranker that fuses a dense embedding list and a sparse lexical list. The method has a small, explicit parameter vector...
By Satyanarayan Pati, Srikanth Patil
Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search describes DocuSearch, an offline multi‑agent system designed for telecom network operations. The system combines semantic vector search, BM25 full‑text search, and knowledge‑graph neighbor expansion, merges the results via Reciprocal Rank Fusion, and reranks with a cross‑encoder before pruning with Maximal Marginal Relevance. A per‑chunk evaluation loop ensures only grounded answers are returned, achieving Precision@10 of 0.69, Recall@10 of 0.79, and an 89.6% grounding rate—improvements of 15, 16, and 18.4 percentage points over a dense‑only baseline.
By Harish Saragadam, Sudhanshu Sharma, Meghana Pujari
MoganColBERT-TR is a Turkish multi‑vector retrieval model that projects token embeddings from 768 to 128 dimensions and uses MaxSim late interaction for scoring. It builds on the previously trained MoganBERT‑TR encoder, adapting it to the ColBERT objective with a single‑epoch distillation phase using cross‑encoder teacher scores. Evaluated on five Turkish BEIR datasets in a zero‑shot setting, it achieves an overall score of 37.36, outperforming larger models on most datasets.
By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay