arXiv AI

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

The paper introduces Difficulty‑Aware Semantic‑ID Optimization (DASO), a post‑training method for generative recommendation that improves tree‑structured item ranking. DASO profiles rollout groups by prefix‑match depth, reallocates a portion of candidates to prefix‑guided completions, and uses a SID‑prefix reward with an auxiliary SFT anchor to address target‑missing failures. On public benchmarks, DASO outperforms MiniOneRec‑style GRPO on 11 of 12 metrics and achieves the best results on 9 of 12 metrics, also improving level‑wise recall on an internal recommendation task.

arXiv AI
Aug 18

Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation

arXiv:2608. 11980v2 Announce Type: replace-cross Abstract: Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autoregressively generating this token sequence.

By Kangning Zhang, Haotian Fang, Xukun Luo, Hao Yin, Yang Gao, Peng Yan, Weiwen Liu, Weinan Zhang, Yong Yu
Hugging Face Trending Papers
6d ago

From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation

The paper introduces a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each autoregressive trace into a history summary, a set of interest hypotheses, and a final SID, a frozen retriever verifies each hypothesis as a catalog query. Rewards are assigned at the hypothesis level when any query retrieves the target within the top‑K, allowing distinct updates for rollouts that share the same SID reward and improving SID recommendation performance on Amazon Reviews datasets.

arXiv AI
6d ago

From Interests to Semantic IDs: Retrieval-Grounded Credit Assignment for Generative Recommendation

The paper proposes a retrieval‑grounded credit‑assignment method for generative recommenders that use Semantic IDs (SIDs). By structuring each generated trace into a history summary, a set of interest hypotheses, and a final SID, and then verifying each hypothesis with a frozen retriever, the method assigns reward at the hypothesis level rather than only at the final SID. Experiments on Amazon Reviews datasets show consistent improvements in SID recommendation, and an oracle analysis on Video Games data demonstrates that selecting target‑relevant queries among generated interests boosts recall and ranking.

By Mengdan Zhu, Yufan Zhao, Yao Zhao, Sophie Di, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao
arXiv AI
6d ago

DeGRe: Dense-supervised Generative Reranking for Recommendation

DeGRe is a dense‑supervised generative reranking framework designed to improve multi‑stage recommender systems by addressing label bias and credit assignment issues. It uses an offline Lookahead Evaluator with beam search to generate dense supervision signals, which are distilled into a lightweight Online Generator that can perform efficient greedy decoding at inference time. Experiments show that DeGRe outperforms baselines on public benchmarks and industrial datasets, and it has been successfully deployed on Taobao Flash Shopping to enhance online recommendations.

By Chaotian Song, Jingyao Zhang, Chenghao Chen, Zisen Sang, Dehai Zhao, Guodong Cao, Boxi Wu, Deng Cai, Jia Jia
arXiv AI
Jul 1

GR2 Technical Report

arXiv:2606. 31984v1 Announce Type: cross Abstract: Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step disproportionately shapes user engagement and downstream performance, particularly for carousel and grid display formats.

By Yufei Li (Yongkang), Zaiwei Zhang (Yongkang), Mingfu Liang (Yongkang), Kavosh Asadi (Yongkang), Jay Xu (Yongkang), Jimmy Kim (Yongkang), Chongyang Bai (Yongkang), Jieyi Zhang (Yongkang), Hongye Xie (Yongkang), Prachi Agrawal (Yongkang), Dian Yu (Yongkang), Tianyi Chen (Yongkang), Jean-Pascal Billaud (Yongkang), Garret Buell (Yongkang), YK (Yongkang), Zhu (Yang), Sachin Patil (Yang), Brooke Bian (Yang), Zhou Fang (Yang), Kevin Huang (Yang), Shiva Sudanagunta (Yang), Yuzhen Huang (Yang), Emma Lu (Yang), Chris O'Brien (Yang), Yang Song (Yang), Lihong Li (Yang), Jacob Tao (Yang), Zhicheng Zhu (Yang), Chao Li (Yang), Gaoxiang Liu (Yang), Neil Wu (Yang), Zhongyin Hu (Yang), Li Han (Yang), Loki Chen (Yang), Ming Lei (Yang), Greg Rehm (Yang), Siyuan Song (Yang), Tianwei Zhang (Yang), Li Li (Yang), Ketan Singh (Yang), Yavuz Yetim (Yang), Ilyas Atishev (Yang), Satendra Gera (Yang), Ashkan Sadeghi (Yang), Rachel Yan (Yang), Nikko Mizutani (Yang), Shuaiwen Wang (Yang), Song Yang (Yang), Zhijing Li (Yang), Jiang Liu (Yang), Mengying Sun (Yang), Fei Tian (Yang), Xiaohan Wei (Yang), Chonglin Sun (Yang), Parish Aggarwal (Yang), Kaushik Rangadurai (Yang), Zhi Hua (Yang), Frank Shyu (Yang), Ruchit Sharma (Yang), Liyuan Li (Yang), Shike Mei (Yang), Wenlin Chen (Yang), Santanu Kolay (Yang), Ben Schulte (Yang), Deepak Chandra (Yang), Adam (Yang), Song, Sandeep Pandey, Xi Liu, Hamed Firooz, Luke Simon
arXiv AI
Jun 9

Generative Reasoning Re-ranker

arXiv:2602. 07774v5 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge.

By Mingfu Liang, Yufei Li, Jay Xu, Kavosh Asadi, Xi Liu, Shuo Gu, Kaushik Rangadurai, Frank Shyu, Shuaiwen Wang, Song Yang, Zhijing Li, Jiang Liu, Mengying Sun, Fei Tian, Xiaohan Wei, Chonglin Sun, Jacob Tao, Shike Mei, Wenlin Chen, Santanu Kolay, Sandeep Pandey, Hamed Firooz, Luke Simon
arXiv AI
6d ago

Learning Better Reasoning for Generative Recommendation with Semantic IDs

The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. It first aligns SIDs with textual and behavioral contexts, then selects candidate reasoning traces that improve ground‑truth item prediction, and finally refines the reasoning policy via reinforcement learning with catalog‑constrained generation and ranking‑aware feedback. Experiments on Amazon Review datasets show Evo‑Rec consistently outperforms existing discriminative, generative, and reasoning‑enhanced recommenders across all metrics.

By Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao