arXiv Machine Learning By Longfeng Wu, Yao Zhou, Tong Zeng, Zhimin Peng, Bhanu Pratap Singh Rawat, Lecheng Zheng, Giovanni Seni, Dawei Zhou

Bi-NAS: Towards Effective and Personalized Explanation for Recommender Systems via Bi-Level Neural Architecture Search

Read the original on arXiv Machine Learning →

arXiv:2607. 01387v1 Announce Type: cross Abstract: Recommender systems are vital in helping users navigate vast amounts of information, offering personalized suggestions and effective explanations for these recommendations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

The Utility of LLMs in Recommender Systems Explanation Evaluation

The paper investigates how large language models (LLMs) can evaluate explanations in recommender systems. It generates 18 explanation prototypes and has 14 LLMs rate them, comparing the results to human ratings from a user study. Findings show that while LLMs mimic human rating patterns and correlate moderately with human judgments, their absolute agreement is low and varies with model size and evaluation design, leading to four practical recommendations for using LLMs in this context.

By Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein
arXiv Machine Learning
Sep 24

A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems

The paper introduces a systematic benchmark for evaluating explainable methods that attribute temporal interactions in sequential recommendation systems. Using a dual-model masking metric, it assesses ten XAI techniques across CNN, Transformer, SASRec, and BERT4Rec backbones on KuaiRand and MovieLens datasets, revealing that gradient-based methods like GradientSHAP and Integrated Gradients are the most faithful and robust. It also finds that raw attention weights are unreliable, while gradient-weighted attention works better on short sequences but degrades on longer horizons, and that faithful methods capture genuine task structure rather than recency or popularity bias.

By Akash Pandey, Kanisha Shah, Addrish Roy, Dwipam Katariya, Hongyangyang Shi, Amanda Ding, Kalanand Mishra, Pranab Mohanty
arXiv Computation and Language
Aug 31

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

The paper investigates whether large language models can learn and reproduce annotator‑specific label‑explanation behavior, using two sentence‑pair tasks with four annotators each. It finds that individual annotator patterns are weak at the single‑annotation level but become detectable after reducing input‑content effects and aggregating across annotators. The authors propose cross‑annotator preference optimization (CAPO), which improves upon prompting and supervised fine‑tuning by better capturing annotator‑specific reasoning while maintaining stable attribution.

By Beiduo Chen, Pingjun Hong, Ziyun Zhang, Benjamin Roth, Anna Korhonen, Barbara Plank