arXiv AI By Noah Mami\'e, Laurin van den Bergh

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

Read the original on arXiv AI →

This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 17

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.

arXiv AI
Sep 23

Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix

The paper introduces a counterfactual observability framework for Netflix’s recommender systems, aiming to disentangle raw engagement signals—such as views and clicks—from confounding factors like content quality, model behavior, presentation bias, and audience reach. It proposes three stakeholder‑centered principles and measurement methods that reduce bias, assess relativity, and capture incrementality, applicable to both single‑stage and cascading recommender architectures. The framework is demonstrated through multiple production deployments, showing its effectiveness in enhancing observability across Netflix’s recommendation pipelines.

By Chaoran Guo, Ding Tong, Ting-Po Lee, Scarlet Chen
arXiv AI
Sep 3

The Utility of LLMs in Recommender Systems Explanation Evaluation

The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.

By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv AI
Sep 2

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

The study evaluates large language models (LLMs) as rerankers in conversational movie recommendation, comparing proprietary, open-weight, and fine-tuned LLMs against collaborative-filtering and sequential baselines on the ReDial benchmark. Results show that the best proprietary LLM achieves an NDCG@10 of 0.1497 with a shared semantic candidate pool, outperforming non-LLM baselines, while open-weight LLMs do not surpass a tuned shallow autoencoder under the same protocol. The analysis also highlights that reranker performance is highly sensitive to candidate generation, pool size, scoring policy, and decoding temperature, suggesting these factors should be reported as standard evaluation fields.

By Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic