arXiv:2608. 11061v1 Announce Type: new Abstract: Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary.
By Artyom Sabitov, Daniil Volkov, Alexey Zaytsev
Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token-level generation task. Despite their promise, we identify a pervasive yet underexplored issue: $\textit{Length Bias}$.
DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.
By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia
arXiv:2607. 19357v1 Announce Type: new Abstract: Recent advances in recommender systems (RS) have shown substantial performance gains through generative modelling.
By Dmitrii Moor, Ben Carterette, Senthilkumar Krishnamoorthy, Kyle Kretschman, Denis Beslic, Melissa Yalla, Alice Y Wang, Mounia Lalmas
The paper introduces rEDMRec, a method that compresses a large language model’s reasoning about user preferences and item comparisons into a compact, editable memory. This memory, organized into four channels—long‑term preference, short‑term context, item perception, and counterfactual hard‑negative comparisons—can be updated by an LLM controller and queried by a lightweight student LLM for ranking, eliminating the need to re‑run the expensive teacher model for each request. Experiments on ML‑1M, Amazon Beauty, and Steam datasets show that rEDMRec consistently outperforms zero‑shot, few‑shot, RAG, and GraphRAG baselines, achieving up to a 13.3% improvement in HR@1 on ML‑1M.
By Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen
arXiv:2403. 00802v2 Announce Type: replace-cross Abstract: Production-grade recommender systems rely heavily on a large-scale corpus used by online media services, including Netflix, Pinterest, and Amazon.
By Amit Kumar Jaiswal
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
arXiv:2607. 04270v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token-level generation task.
By Hongchen Li, Bohao Wang, Jingbang Chen, Weiqin Yang, Hang Pan, Bingde Hu, Can Wang, Jiawei Chen
arXiv:2608.24079v1 Announce Type: cross
Abstract: A shared search-and-recommendation index must score new items from features alone because search has no exploration slot. In a public log covering bo...
By Theodore Rogers, Joe Standerfer, Dmitrii Timoshenko, Haoxue Li, Zuhaib Akhtar, Soyoung Yang
The paper investigates the trade‑off of using a shared search‑and‑recommendation index that scores new items purely from features, thereby keeping the index open to unseen items. Experiments on public logs show that a feature‑based tower can match warm‑item performance (Recall@20 0.9595 vs 0.9510) and a lexical baseline, while a full‑catalog check is inconclusive. The study also quantifies the cost of this openness on recommendation quality across several baselines, revealing that exact full‑softmax training improves recall but is impractical at catalog scale.
By Theodore Rogers, Joe Standerfer, Dmitrii Timoshenko, Haoxue Li, Zuhaib Akhtar, Soyoung Yang
arXiv:2603.02561v2 Announce Type: replace-cross
Abstract: Attention mechanism remains the defining operator in Transformers since it provides expressive global credit assignment, yet its quadratic co...
By Chenghao Zhang, Chao Feng, Yuanhao Pu, Xunyong Yang, Wenhui Yu, Xiang Li, Chunjie Chen, Kaiqiao Zhan
The paper examines LLM-based recommendation rerankers that are often evaluated under an oracle protocol, which guarantees the ground-truth item is present in the scored set. Across Amazon datasets, this protocol overestimates realistic NDCG@10 by 92–95% because realistic retrieval only covers 2–19% of relevant items at K=100, creating a recall ceiling that limits any closed-candidate reranker's top‑k NDCG. The authors find that various optimisation strategies—including prompt engineering, model scaling, sequential models, supervised neural rerankers, LoRA fine‑tuning, hybrid retrieval, score‑aware prompting, and LLM+CF fusion—do not significantly improve over a collaborative‑filtering baseline under realistic retrieval, and they propose a Recall‑Aware Evaluation Protocol (RAEP) to better assess rerankers in low‑recall regimes.
By Zhaohui Wang