Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token-level generation task. Despite their promise, we identify a pervasive yet underexplored issue: $\textit{Length Bias}$.
arXiv:2607. 04270v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token-level generation task.
By Hongchen Li, Bohao Wang, Jingbang Chen, Weiqin Yang, Hang Pan, Bingde Hu, Can Wang, Jiawei Chen
The paper introduces rEDMRec, a method that compresses a large language model’s reasoning about user preferences and item comparisons into a compact, editable memory. This memory, organized into four channels—long‑term preference, short‑term context, item perception, and counterfactual hard‑negative comparisons—can be updated by an LLM controller and queried by a lightweight student LLM for ranking, eliminating the need to re‑run the expensive teacher model for each request. Experiments on ML‑1M, Amazon Beauty, and Steam datasets show that rEDMRec consistently outperforms zero‑shot, few‑shot, RAG, and GraphRAG baselines, achieving up to a 13.3% improvement in HR@1 on ML‑1M.
By Minh Hoang Nguyen, Tung Le, Huy Tien Nguyen
RPCBench is a new benchmark designed to evaluate large language models’ ability to critique recommendation requests by detecting, diagnosing, and handling flawed premises. It includes evidence‑grounded test instances across five recommendation domains and ten types of premise failures, and introduces a fine‑grained evaluation framework covering detection, error localization, handling strategy, and evidence faithfulness. Experiments with 11 LLMs reveal that proactive detection is the main bottleneck, with models struggling most on underspecified‑premise errors and showing that optimal critique quality occurs at intermediate reasoning lengths.
By Zhongru Chen, Yuan Wu, Yi Chang
The paper introduces CoVeMem, a Collaborative Vector Memory system that replaces text-based memory in agentic recommender systems with vectorized user and item states derived from a frozen LightGCN model. By retrieving relevant historical states at each decision and integrating them as soft tokens in the LLM’s context, CoVeMem enables contrastive alignment and listwise co‑training to learn how to read and rank these states, achieving performance on par with or better than existing text‑memory agents across multiple benchmarks without extra LLM calls for memory updates.
By Hanchong Chen, Xing Tang, Lingjie Li, Xiongfeng Shan, Xiuqiang He
The paper presents a method for fine‑tuning a large language model (LLM) recommender to generate personalized, non‑harmful explanations for its recommendations. By training two LLM‑judge reward models and using constrained GRPO, the authors achieve a significant increase in the PASS rate for all three criteria, from 0.649 to 0.956 on their own judges and from 0.677 to 0.931 on an independent judge. The fine‑tuned model maintains its original recommendation performance, demonstrating that LLM‑based recommenders can be adapted to complex tasks without loss of effectiveness.
By Jiashu He, Emma Yanyang Kong, JJ Tan, David Fagnan
The paper proposes an incremental recommendation system that uses causal modeling to avoid delivering redundant recommendations. By leveraging existing holdback data and a dual‑threshold targeting policy, the authors reduce recommendation impressions by 7% without harming overall content consumption. Joint training with holdback data also improves the calibration of the treated model, suggesting better generalisable representations than purely observational models.
By Athanasios Vlontzos, David Gustafsson, Michael O'Riordan, Ciar\'an M. Gilligan-Lee
This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.
By Noah Mami\'e, Laurin van den Bergh
The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.
By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
The paper proposes an incremental recommendation approach that uses a causal model built from existing holdback data to avoid delivering redundant recommendations. By applying a dual‑threshold targeting policy, the system only recommends content when the likelihood of a treated stream is high and the likelihood of an organic stream is low, thereby reducing recommendation impressions by 7% without hurting overall consumption. Joint training with holdback data also improves the calibration of the treated head, suggesting that causal models capture more generalisable representations than purely observational models.
The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.
arXiv:2606. 00467v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions.
By Etienne Casanova, Rafal Kocielnik, R. Michael Alvarez