DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.
By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia
arXiv:2511.22707v2 Announce Type: replace-cross
Abstract: In web environments, user preferences are often refined progressively as users move from browsing broad categories to exploring specific item...
By Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua, Jingrui He
The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. It first aligns SIDs with textual and behavioral contexts, then selects candidate reasoning traces that improve ground‑truth item prediction, and finally refines the reasoning policy via reinforcement learning with catalog‑constrained generation and ranking‑aware feedback. Experiments on Amazon Review datasets show Evo‑Rec consistently outperforms existing discriminative, generative, and reasoning‑enhanced recommenders across all metrics.
By Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao
The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. First, it aligns SIDs with textual and behavioral contexts; second, it samples and selects reasoning traces that improve ground‑truth item prediction via supervised fine‑tuning; third, it refines the reasoning policy with reinforcement learning using catalog‑constrained generation and ranking‑aware feedback. Experiments on three Amazon Review datasets show Evo‑Rec consistently outperforms discriminative, generative, and other reasoning‑enhanced recommenders across all metrics.
arXiv:2608. 11061v1 Announce Type: new Abstract: Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary.
By Artyom Sabitov, Daniil Volkov, Alexey Zaytsev
arXiv:2609.15094v1 Announce Type: cross
Abstract: In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user pop...
By Yi Chen, Rufeng Cheng, Qiang Xie, Tao Li
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items $K$, the final classification layer dominates memory, requiring $O(nK)$ logits and gradients to materialize for a batch of $n$ examples.
arXiv:2608. 14011v1 Announce Type: cross Abstract: Generative recommendation autoregressively generates the semantic IDs of the target item, unifying preference modeling and index retrieval within the shared token space.
By Haokai Ma, Aoqi Hu, Yueao Xing, Ruobing Xie, Yonghui Yang, Teng Tu, Lei Meng, Tat-Seng Chua
arXiv:2506. 16114v3 Announce Type: replace-cross Abstract: Generative recommendations (GR), which usually include item tokenizers and generative Large Language Models (LLMs), have demonstrated remarkable success across a wide range of scenarios.
By Yejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu, Xinhang Li, Wenlin Zhang, Feng Li, Pengjie Wang, Chuan Yu, Jian Xu, Bo Zheng, Xiangyu Zhao
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
arXiv:2606. 11023v1 Announce Type: cross Abstract: Sequential recommendation aims to predict users' next interaction with items by analyzing their historical behavior.
By Yifan Li, Jiahong Liu, Xinni Zhang, Hao Chen, Yankai Chen, Wenhao Yu, Jianting Chen, Irwin King
arXiv:2606. 01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy.
By Wyame Benslimane, Tinghan Ye, Pascal Van Hentenryck, Paul Grigas