The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. It first aligns SIDs with textual and behavioral contexts, then selects candidate reasoning traces that improve ground‑truth item prediction, and finally refines the reasoning policy via reinforcement learning with catalog‑constrained generation and ranking‑aware feedback. Experiments on Amazon Review datasets show Evo‑Rec consistently outperforms existing discriminative, generative, and reasoning‑enhanced recommenders across all metrics.
By Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao
arXiv:2608. 11980v2 Announce Type: replace-cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence.
By Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
arXiv:2608. 11980v1 Announce Type: cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence.
By Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. First, it aligns SIDs with textual and behavioral contexts; second, it samples and selects reasoning traces that improve ground‑truth item prediction via supervised fine‑tuning; third, it refines the reasoning policy with reinforcement learning using catalog‑constrained generation and ranking‑aware feedback. Experiments on three Amazon Review datasets show Evo‑Rec consistently outperforms discriminative, generative, and other reasoning‑enhanced recommenders across all metrics.
arXiv:2603. 23183v2 Announce Type: replace-cross Abstract: Recent advances in generative recommendation have leveraged pretrained LLMs by formulating sequential recommendation as autoregressive generation over a unified token space comprising language tokens and itemic identifiers, where each item is represented by a compact sequence of discrete tokens, namely Semantic IDs (SIDs).
By Yingzhi He, Yan Sun, Junfei Tan, Yuxin Chen, Xiaoyu Kong, Chunxu Shen, Xiang Wang, An Zhang, Tat-Seng Chua
The paper introduces Difficulty‑Aware Semantic‑ID Optimization (DASO), a post‑training method for generative recommendation that improves tree‑structured item ranking. DASO profiles rollout groups by prefix‑match depth, reallocates a portion of candidates to prefix‑guided completions, and uses a SID‑prefix reward with an auxiliary SFT anchor to address target‑missing failures. On public benchmarks, DASO outperforms MiniOneRec‑style GRPO on 11 of 12 metrics and achieves the best results on 9 of 12 metrics, also improving level‑wise recall on an internal recommendation task.
By Xin Yu, Stephen Li, Sina Aghaei, Zifan Zhu, Jiamu Bai, Guanjie Huang, Bo Peng, Yiyao Liu, Lingzhou Xue
The paper introduces a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each autoregressive trace into a history summary, a set of interest hypotheses, and a final SID, a frozen retriever verifies each hypothesis as a catalog query. Rewards are assigned at the hypothesis level when any query retrieves the target within the top‑K, allowing distinct updates for rollouts that share the same SID reward and improving SID recommendation performance on Amazon Reviews datasets.
arXiv:2602. 07774v5 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge.
By Mingfu Liang, Yufei Li, Jay Xu, Kavosh Asadi, Xi Liu, Shuo Gu, Kaushik Rangadurai, Frank Shyu, Shuaiwen Wang, Song Yang, Zhijing Li, Jiang Liu, Mengying Sun, Fei Tian, Xiaohan Wei, Chonglin Sun, Jacob Tao, Shike Mei, Wenlin Chen, Santanu Kolay, Sandeep Pandey, Hamed Firooz, Luke Simon
The paper proposes a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each generated trace into a history summary, a set of interest hypotheses, and a final SID, and then verifying each hypothesis with a frozen retriever, the method assigns reward at the hypothesis level rather than only at the final SID. Experiments on Amazon Reviews datasets show consistent improvements in SID recommendation, and an oracle analysis on Video Games data demonstrates that selecting target‑relevant queries among generated interests boosts recall and ranking.
By Mengdan Zhu, Yufan Zhao, Yao Zhao, Sophie Di, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao
arXiv:2609.40360v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions...
By Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang
arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.
By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su
arXiv:2606. 14142v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world knowledge.
By Yinhan He, Liam Collins, Bhuvesh Kumar, Jundong Li, Neil Shah, Donald Loveland