Hugging Face Trending Papers

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

The study reproduces a prior work on recommender systems that use generated natural‑language user profiles to enhance transparency and user control. It confirms that the User Profile Recommendation (UPR) model performs competitively and that altering these profiles uniformly shifts predicted ratings without changing ranking order. Additional experiments include context ablation, multi‑seed stability, and mechanistic interpretability analysis with the nnsight framework.

arXiv AI
Sep 18

Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

This reproducibility study confirms that incorporating generated natural‑language user profiles into recommender systems enhances transparency and allows users to directly intervene by correcting preferences or addressing cold‑start issues. The authors replicated the original findings and extended the evaluation with context ablation, multi‑seed stability tests, and mechanistic interpretability analysis using nnsight. Their results show that while perturbing profiles shifts predicted ratings uniformly across genres, the overall rankings remain unchanged, attributing this to the rating‑regression objective rather than the profile interface.

By Noah Mami\'e, Laurin van den Bergh
arXiv AI
Sep 23

Beyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix

The paper introduces a counterfactual observability framework for Netflix’s recommender systems, aiming to disentangle raw engagement signals—such as views and clicks—from confounding factors like content quality, model behavior, presentation bias, and audience reach. It proposes three stakeholder‑centered principles and measurement methods that reduce bias, assess relativity, and capture incrementality, applicable to both single‑stage and cascading recommender architectures. The framework is demonstrated through multiple production deployments, showing its effectiveness in enhancing observability across Netflix’s recommendation pipelines.

By Chaoran Guo, Ding Tong, Ting-Po Lee, Scarlet Chen
arXiv AI
Sep 3

The Utility of LLMs in Recommender Systems Explanation Evaluation

The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.

By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv AI
Sep 10

Neutralizing Popularity Bias in LLM-based Recommendation via Counterfactual Reasoning Guidelines

The paper introduces NPRec, a model‑agnostic framework that uses counterfactual reasoning to neutralize popularity bias in large language model–based recommender systems. By generating debiased textual guidelines that separate intrinsic user interests from popularity signals, NPRec injects these guidelines at inference time to guide the LLM’s generation without updating parameters. Experiments on three real‑world datasets show improved recommendation accuracy, explanation quality, and debiasing performance.

By Guanrong Li, Haolin Yang, Xinyu Liu, Zhen Wu, Rui Xia, Xinyu Dai
arXiv AI
Sep 2

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

The study evaluates large language models (LLMs) as rerankers in conversational movie recommendation, comparing proprietary, open-weight, and fine-tuned LLMs against collaborative-filtering and sequential baselines on the ReDial benchmark. Results show that the best proprietary LLM achieves an NDCG@10 of 0.1497 with a shared semantic candidate pool, outperforming non-LLM baselines, while open-weight LLMs do not surpass a tuned shallow autoencoder under the same protocol. The analysis also highlights that reranker performance is highly sensitive to candidate generation, pool size, scoring policy, and decoding temperature, suggesting these factors should be reported as standard evaluation fields.

By Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic
arXiv AI
Sep 18

FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System

FacetCRS is a conversational recommender system that tackles the filter‑bubble problem by learning multi‑faceted user preferences—entity, word, context, and review facets—through natural language interactions. The framework adaptively models these preference facets and incorporates external knowledge to provide diverse recommendations. Experiments on two benchmark datasets show that FacetCRS outperforms existing methods in reducing filter bubbles and improving recommendation quality.

By Yongsen Zheng, Ziliang Chen, Jinghui Qin, Liang Lin
arXiv Machine Learning
Sep 17

LIGE-GR: A Smooth Leap from Ranking to Generative Recommendation in the LLM Era

LIGE‑GR is a framework that transitions traditional ranking‑based recommender systems to a generative, listwise approach inspired by large language models. It extends existing pointwise recommendation models into a listwise generation system, enabling sequence‑level optimization without overhauling the entire infrastructure. Experiments on Instagram Reels and Facebook Video show modest gains in user time spent—1.14 % and 0.72 % respectively—while adding only slight inference overhead.

By Venkat Srinivas, Chenzhang He, Sam Woodmansee, Shawn Lian, Wenjie Hu, Renjie Jiang, Ziheng Huang, Xinyuan Zhang, Zhihao Zheng, Zhuoran Yu, Rui Li, Lei Yuan, Ziwei Li, Jimmy Jia, Mert Terzihan, Ekrem Kocaguneli, Yiming Liao, Zhichen Zhao, Yue Yin, Yue Weng, Wanlin Ma, Xufeng Cai, Weimiao Wu, Yezhou Huang, Du Zhang, Yukun Ding, Aaron Johnston, Yueming Wang, Zhaojie Gong, Yuting Zhang, Serena Li, Adithya Ganesh, Boying Liu, Haichuan Yang, Xialu Li, Matt Ma, Qunshu Zhang, John Joshua Miller, Praveen Rathinavelu, Cheng Huang, Aadhar Sachdeva, Josh Karns, Andres Aaron Gutierrez, Neil Agarwal, Gustas Pladis, Vladimir Batygin, Gopal Ray, Aditya Priyadarshi, Shantanu Patil, Zhe Wang, Penny Pan, Yiping Han, Arun Singh, Guangdeng Liao, Bi Xue, Xinyao Hu, Yang Song, Yisong Song, Meihong Wang, Haotian Wu, Deepak Agarwal, Ji Liu