arXiv Machine Learning

Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation

arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.

arXiv Machine Learning
Sep 24

The Recall Ceiling of LLM Recommendation Reranking

The paper examines LLM-based recommendation rerankers that are often evaluated under an oracle protocol, which guarantees the ground-truth item is present in the scored set. Across Amazon datasets, this protocol overestimates realistic NDCG@10 by 92–95% because realistic retrieval only covers 2–19% of relevant items at K=100, creating a recall ceiling that limits any closed-candidate reranker's top‑k NDCG. The authors find that various optimisation strategies—including prompt engineering, model scaling, sequential models, supervised neural rerankers, LoRA fine‑tuning, hybrid retrieval, score‑aware prompting, and LLM+CF fusion—do not significantly improve over a collaborative‑filtering baseline under realistic retrieval, and they propose a Recall‑Aware Evaluation Protocol (RAEP) to better assess rerankers in low‑recall regimes.

By Zhaohui Wang
arXiv AI
Sep 2

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

The study evaluates large language models (LLMs) as rerankers in conversational movie recommendation, comparing proprietary, open-weight, and fine-tuned LLMs against collaborative-filtering and sequential baselines on the ReDial benchmark. Results show that the best proprietary LLM achieves an NDCG@10 of 0.1497 with a shared semantic candidate pool, outperforming non-LLM baselines, while open-weight LLMs do not surpass a tuned shallow autoencoder under the same protocol. The analysis also highlights that reranker performance is highly sensitive to candidate generation, pool size, scoring policy, and decoding temperature, suggesting these factors should be reported as standard evaluation fields.

By Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic
arXiv Machine Learning
Jul 14

Serving the Long Tail: Training-Free LLM Candidate Generation for Vacation Rental Marketplaces

arXiv:2607. 09877v1 Announce Type: new Abstract: Vacation rental marketplaces face a structural imbalance on the supply side: a small fraction of properties receive most user interactions, while the long tail of new, niche, and seasonal listings generates too little behavioral signal for collaborative filtering to serve effectively.

By Syed Mohammed Arshad Zaidi, Eric Rincon, Shayan Hassantabar
arXiv Computation and Language
Sep 23

ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap

The paper reports the ABAI submission to COLIEE 2026 Task 1, a case law retrieval challenge that suppresses cited passages, and details a four‑stage retrieval pipeline: multi‑view BM25 with reciprocal rank fusion, neural reranking, graph‑based features via a graph attention network, and a LightGBM meta‑learner over 34 features. The best run achieved an F1 score of 0.177 on the official test set, compared to a cross‑validated 0.311, and the authors attribute the gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. A controlled post‑hoc study examined the impact of threshold transfer, decision quality across time, and query similarity, and identified specific remedies—such as BM25 length‑normalisation tuning, event‑triple views, and dense fusion—that improved recall, while other interventions had no effect.

By Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han
arXiv Machine Learning
Aug 28

Keeping the Index Open: The Recommendation-Side Cost of Shared Search and Recommendation

The paper investigates the trade‑off of using a shared search‑and‑recommendation index that scores new items purely from features, thereby keeping the index open to unseen items. Experiments on public logs show that a feature‑based tower can match warm‑item performance (Recall@20 0.9595 vs 0.9510) and a lexical baseline, while a full‑catalog check is inconclusive. The study also quantifies the cost of this openness on recommendation quality across several baselines, revealing that exact full‑softmax training improves recall but is impractical at catalog scale.

By Theodore Rogers, Joe Standerfer, Dmitrii Timoshenko, Haoxue Li, Zuhaib Akhtar, Soyoung Yang
arXiv Computation and Language
Aug 27

Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting

The paper introduces “PACE”, a training‑free framework that tackles bottlenecks in Retrieval‑Augmented Generation by frontloading evidence and adaptively budgeting reranking. It first reorders candidate documents based on marginal evidence coverage—prioritizing query‑relevant, complementary, and chain‑forming documents—providing a $(1-1/e)$ approximation guarantee. Then it dynamically adjusts the reranking budget according to the relative pressure of the reranker and the language model, improving evidence recall and reducing p95 latency in multi‑hop QA workloads.

By Weibin Cai, Reza Zafarani
arXiv AI
Aug 13

CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation

arXiv:2605. 04495v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing retrieval and reranking methods primarily estimate query--document relevance.

By Zhipeng Song, Yizhi Zhou, Xiangyu Kong, Jiulong Jiao, Xuezhou Ye, Chunqi Gao, Xueqing Shi, Yu Wang, Yuhang Zhou, Heng Qi