arXiv AI

Reasoning with Evidence, Not Merely Rationales: Verifiable Preference Proofs for LLM-Based Recommendation

The paper introduces PROVE-REC, a two‑pass framework that generates verifiable preference proofs for large language model (LLM) recommendation systems. In Pass A, the model condenses user interaction histories into a compact proof of positive and avoidance claims linked to specific evidence entries. Pass B then uses only this proof and its evidence to predict the next item, ensuring the recommendation follows the reasoning path. Verification steps compare masked evidence and removed claims to confirm grounding and influence, while a ranking‑preservation objective retains useful historical information. Experiments on diverse real‑world datasets show PROVE‑REC outperforms strong baselines by up to 7.45%, producing claims that are both better grounded and more influential to recommendation quality.

arXiv AI
Sep 15

Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself

The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.

By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
arXiv AI
Sep 2

RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

RPCBench is a new benchmark designed to evaluate large language models’ ability to critique recommendation requests by detecting, diagnosing, and handling flawed premises. It includes evidence‑grounded test instances across five recommendation domains and ten types of premise failures, and introduces a fine‑grained evaluation framework covering detection, error localization, handling strategy, and evidence faithfulness. Experiments with 11 LLMs reveal that proactive detection is the main bottleneck, with models struggling most on underspecified‑premise errors and showing that optimal critique quality occurs at intermediate reasoning lengths.

By Zhongru Chen, Yuan Wu, Yi Chang
arXiv AI
Sep 10

Neutralizing Popularity Bias in LLM-based Recommendation via Counterfactual Reasoning Guidelines

The paper introduces NPRec, a model‑agnostic framework that uses counterfactual reasoning to neutralize popularity bias in large language model–based recommender systems. By generating debiased textual guidelines that separate intrinsic user interests from popularity signals, NPRec injects these guidelines at inference time to guide the LLM’s generation without updating parameters. Experiments on three real‑world datasets show improved recommendation accuracy, explanation quality, and debiasing performance.

By Guanrong Li, Haolin Yang, Xinyu Liu, Zhen Wu, Rui Xia, Xinyu Dai
arXiv AI
Sep 25

Learning Better Reasoning for Generative Recommendation with Semantic IDs

The paper introduces Evo-Rec, a three‑stage framework that improves generative recommendation by learning better reasoning traces for Semantic ID (SID) generation. It first aligns SIDs with textual and behavioral contexts, then selects candidate reasoning traces that improve ground‑truth item prediction, and finally refines the reasoning policy via reinforcement learning with catalog‑constrained generation and ranking‑aware feedback. Experiments on Amazon Review datasets show Evo‑Rec consistently outperforms existing discriminative, generative, and reasoning‑enhanced recommenders across all metrics.

By Mengdan Zhu, Yufan Zhao, Sophie Di, Yao Zhao, Tao Di, Yulan Yan, Sridhar Iyer, Liang Zhao
arXiv AI
Sep 3

The Utility of LLMs in Recommender Systems Explanation Evaluation

The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.

By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv AI
Jun 9

Generative Reasoning Re-ranker

arXiv:2602. 07774v5 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge.

By Mingfu Liang, Yufei Li, Jay Xu, Kavosh Asadi, Xi Liu, Shuo Gu, Kaushik Rangadurai, Frank Shyu, Shuaiwen Wang, Song Yang, Zhijing Li, Jiang Liu, Mengying Sun, Fei Tian, Xiaohan Wei, Chonglin Sun, Jacob Tao, Shike Mei, Wenlin Chen, Santanu Kolay, Sandeep Pandey, Hamed Firooz, Luke Simon
arXiv AI
Sep 10

Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation

The paper introduces Fact-Ablated Evaluation (FAE), a framework that iteratively removes cited evidence to test whether large language models (LLMs) adjust their fact‑checking predictions accordingly. Experiments reveal that many off‑the‑shelf LLMs rely more on internal knowledge than on the provided evidence. To address this, the authors propose REAL, a training method that uses counterfactual evidence supervision to encourage LLMs to base veracity judgments on evidence, achieving better evidence‑dependent performance across four datasets.

By Xingyu Deng, Mingzi Cao, Nikolaos Aletras, Xi Wang, Mark Stevenson
arXiv Machine Learning
Jul 14

RecRec: Recursive Refinement for Sequential Recommendation

arXiv:2607. 10541v1 Announce Type: cross Abstract: Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refinement, relying on increasingly deep architectures to capture complex patterns.

By Pervez Shaik, Prosenjit Biswas, Abhinav Thorat, Ravi Kolla, Niranjan Pedanekar
arXiv AI
Jun 2

Principled Synthetic Data Enables the First Scaling Laws for LLMs in Recommendation

arXiv:2602. 07298v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) represent a promising frontier for recommender systems, yet their development has been impeded by the absence of predictable scaling laws, which are crucial for guiding research and optimizing resource allocation.

By Benyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen, Qifei Wang, Wei Sun, Shen Li, Jia Li, Jiahao Wu, Qunshu Zhang, Neeraj Bhatia, Xiangjun Fan, Hong Yan
arXiv Machine Learning
Sep 17

LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era

LIGE‑GR is a framework that transitions traditional ranking‑based recommender systems to a generative, listwise approach inspired by large language models. It extends existing pointwise recommendation models into a listwise generation system, enabling sequence‑level optimization without overhauling the entire infrastructure. Experiments on Instagram Reels and Facebook Video show modest gains in user time spent—1.14 % and 0.72 % respectively—while adding only slight inference overhead.

By Venkat Srinivas, Chenzhang He, Sam Woodmansee, Shawn Lian, Wenjie Hu, Renjie Jiang, Ziheng Huang, Xinyuan Zhang, Zhihao Zheng, Zhuoran Yu, Rui Li, Lei Yuan, Ziwei Li, Jimmy Jia, Mert Terzihan, Ekrem Kocaguneli, Yiming Liao, Zhichen Zhao, Yue Yin, Yue Weng, Wanlin Ma, Xufeng Cai, Weimiao Wu, Yezhou Huang, Du Zhang, Yukun Ding, Aaron Johnston, Yueming Wang, Zhaojie Gong, Yuting Zhang, Serena Li, Adithya Ganesh, Boying Liu, Haichuan Yang, Xialu Li, Matt Ma, Qunshu Zhang, John Joshua Miller, Praveen Rathinavelu, Cheng Huang, Aadhar Sachdeva, Josh Karns, Andres Aaron Gutierrez, Neil Agarwal, Gustas Pladis, Vladimir Batygin, Gopal Ray, Aditya Priyadarshi, Shantanu Patil, Zhe Wang, Penny Pan, Yiping Han, Arun Singh, Guangdeng Liao, Bi Xue, Xinyao Hu, Yang Song, Yisong Song, Meihong Wang, Haotian Wu, Deepak Agarwal, Ji Liu