The paper examines LLM-based recommendation rerankers that are often evaluated under an oracle protocol, which guarantees the ground-truth item is present in the scored set. Across Amazon datasets, this protocol overestimates realistic NDCG@10 by 92–95% because realistic retrieval only covers 2–19% of relevant items at K=100, creating a recall ceiling that limits any closed-candidate reranker's top‑k NDCG. The authors find that various optimisation strategies—including prompt engineering, model scaling, sequential models, supervised neural rerankers, LoRA fine‑tuning, hybrid retrieval, score‑aware prompting, and LLM+CF fusion—do not significantly improve over a collaborative‑filtering baseline under realistic retrieval, and they propose a Recall‑Aware Evaluation Protocol (RAEP) to better assess rerankers in low‑recall regimes.
By Zhaohui Wang
arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.
By Zhe Dong (University of Maine at Presque Isle), Fang Qin (Stanford University), Manish Shah (Independent Researcher), Yicheng Wang (Independent Researcher)
TailSpec-EASE is a lightweight linear recommender that incorporates a relation‑aware spectral knowledge‑graph prior into a local closed‑form reconstruction objective. By adapting the prior strength to item popularity, it provides stronger semantic guidance for long‑tail items. Across four public benchmarks, it achieves a favorable balance of overall accuracy, long‑tail performance, and training cost, improving NDCG@20 by up to 24% over a no‑KG baseline and training in just 37 seconds on CPU compared to thousands of seconds for GPU‑based KGAT and CPU LightGCN.
By Jianru Shen
arXiv:2506. 07449v2 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLMs) have driven their adoption in recommender systems through Retrieval-Augmented Generation (RAG) frameworks.
By Vahid Azizi, Fatemeh Koochaki
The paper introduces a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each autoregressive trace into a history summary, a set of interest hypotheses, and a final SID, a frozen retriever verifies each hypothesis as a catalog query. Rewards are assigned at the hypothesis level when any query retrieves the target within the top‑K, allowing distinct updates for rollouts that share the same SID reward and improving SID recommendation performance on Amazon Reviews datasets.
arXiv:2602. 07774v5 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge.
By Mingfu Liang, Yufei Li, Jay Xu, Kavosh Asadi, Xi Liu, Shuo Gu, Kaushik Rangadurai, Frank Shyu, Shuaiwen Wang, Song Yang, Zhijing Li, Jiang Liu, Mengying Sun, Fei Tian, Xiaohan Wei, Chonglin Sun, Jacob Tao, Shike Mei, Wenlin Chen, Santanu Kolay, Sandeep Pandey, Hamed Firooz, Luke Simon