Recommender systems are vital in helping users navigate vast amounts of information, offering personalized suggestions and effective explanations for these recommendations. While previous efforts have attempted to provide such explanations, evaluating their effectiveness across various scenarios remains a challenge.
The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.
By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv:2606. 18897v1 Announce Type: cross Abstract: Intent-based recommender systems have gained significant attention for improving accuracy and interpretability by modeling the underlying motivations behind user behaviors.
By Jiangnan Xia, Xuansheng Wu, Yu Yang, Xin Wang, Ninghao Liu
The paper introduces a systematic benchmark for evaluating explainable methods that attribute temporal interactions in sequential recommendation systems. Using a dual-model masking metric, it assesses ten XAI techniques across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens datasets, revealing that gradient-based methods like GradientSHAP and Integrated Gradients are the most faithful and robust. It also finds that raw attention weights are unreliable, while gradient-weighted attention works better on short sequences but degrades on longer horizons, and that faithful methods capture genuine task structure rather than recency or popularity bias.
By Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
arXiv:2606. 08497v1 Announce Type: new Abstract: As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability.
By Minyoung Hwang, Seokhyun Lee, Changhee Lee
The paper investigates whether large language models can learn and reproduce annotator‑specific label‑explanation behavior, using two sentence‑pair tasks with four annotators each. It finds that individual annotator patterns are weak at the single‑annotation level but become detectable after reducing input‑content effects and aggregating across annotators. The authors propose cross‑annotator preference optimization (CAPO), which improves upon prompting and supervised fine‑tuning by better capturing annotator‑specific reasoning while maintaining stable attribution.
By Beiduo Chen, Pingjun Hong, Ziyun Zhang, Benjamin Roth, Anna Korhonen, Barbara Plank
As artificial intelligence and machine learning (AI/ML) models become integral to network operations, their lack of transparency poses a significant barrier to operator trust. Existing explainable artificial intelligence (XAI) techniques often fail to bridge this gap for non-specialists, producing technical outputs that are difficult to translate into actionable insights.
AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices.
arXiv:2603. 02070v3 Announce Type: replace Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise.
By Guilhem Fouilh\'e, Rebecca Eifler, Antonin Poch\'e, Sylvie Thi\'ebaux, Nicholas Asher
arXiv:2602. 12612v2 Announce Type: replace-cross Abstract: Traditional methods for automating recommender system design, such as Neural Architecture Search (NAS), are often constrained by a fixed search space defined by human priors, limiting innovation to pre-defined operators.
By Sein Kim, Sangwu Park, Hongseok Kang, Wonjoong Kim, Jimin Seo, Yeonjun In, Kanghoon Yoon, Hyunsik Jeon, Chanyoung Park
Re2A is a new framework for situated conversational recommendation that models user interactions within shared physical environments. It introduces rubric-based preference reasoning to explicitly capture user preferences from dialogue history and scene context, and a preference-conditioned optimization to align generated responses with both user satisfaction and situational consistency. Experiments on two SCR datasets show that Re2A outperforms existing methods, providing more precise and context-aware recommendations.
By Dongding Lin, Jian Wang, Xiaoyan Zhao, Wenjie Li
arXiv:2608.20801v1 Announce Type: cross
Abstract: While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains chal...
By Dojun Hwang, Seunghan Lee, Cheonyoung Park, Sara Yu, SeongKu Kang